Evidence Ledgersource .md4 sections

Evidence Ledger

AI Interpretability Paper | Foote and Pink | Compiled August 18, 2026

Every entry below was checked against a primary source. Where the figure in an earlier working document did not survive that check, the discrepancy is flagged in bold. Entries are grouped by the job each source does in the argument, because the same study is often load-bearing for a different reason in each of the four candidate spines.


1 of 4. Tier 1: The empirical spineThese are the studies with numbers attached.

These are the studies with numbers attached. They are what makes the argument publishable rather than essayistic.

Budzyń, K. et al. (2025)

"Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study." The Lancet Gastroenterology & Hepatology. DOI: 10.1016/S2468-1253(25)00133-5

Nineteen experienced endoscopists across four Polish centres, each with more than 2,000 prior procedures. Adenoma detection rate on non-AI-assisted colonoscopies fell from 28.4% (226/795) in the three months before routine AI introduction to 22.4% (145/648) in the period after. A 6 percentage point absolute decline, roughly 20% relative.

Why it matters: this is Bainbridge's 1983 irony measured in live clinical practice among credentialed experts, published in a Lancet title. It is the single strongest answer to "isn't this just moral panic." It also moves the argument out of the classroom, which matters because education-only evidence invites the reply that students are a special case.

Caveats to state yourself: observational, not randomized. Confounds with time period and case mix are plausible. An accompanying comment in the same journal (Zhou, 2025) discusses exactly these.

Bansal, G., Wu, T., Zhou, J. et al. (2021)

"Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance." CHI 2021. DOI: 10.1145/3411764.3445717

1,626 participants across three datasets (beer reviews, Amazon book reviews, LSAT logical reasoning). Verbatim finding: explanations "increased the chance that humans will accept the AI's recommendation, regardless of its correctness."

Why it matters: the reflexive turn. Interpretability output is itself a persuasive medium. This is what licenses applying media theory to interpretability artifacts rather than only to model outputs.

Do not cite this alone. See Vasconcelos below.

Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M. et al. (2023)

"Explanations Can Reduce Overreliance on AI Systems During Decision-Making." Proc. ACM Hum.-Comput. Interact. 7, CSCW1. arXiv:2212.06823

Five studies, 731 participants, maze-solving task with a simulated AI. Explanations do reduce overreliance, conditional on a cost-benefit calculation: people engage with an explanation when doing so costs less effort than verifying the task themselves. The authors read prior null results as cases where the explanation failed to lower verification cost.

This is the most important correction to the current argument. The clean claim "explanations manufacture compliance" is no longer supportable as stated. The upgraded claim is stronger and more useful:

An explanation only changes behavior when reading it is cheaper than checking the work.

That is a design principle, a pedagogical principle, and a testable prediction. It also converts the paper's stance from critique into construction, which is the register that gets cited.

Vaccaro, M., Almaatouq, A. and Malone, T. (2024)

"When combinations of humans and AI are useful: a systematic review and meta-analysis." Nature Human Behaviour 8, 2293-2303. arXiv:2405.06087

370 unique effect sizes from 106 experiments published January 2020 to June 2023 (5,126 papers screened, 74 met inclusion criteria).

Correction to transparency_problem_map_1.html: the map states "Gains appeared in creation tasks (g = +0.19)." The confidence interval crosses zero. That result is directional at best and must not be reported as a finding. The overall and decision-task effects are solid and are the ones to lead with.

Why it matters: pre-empts the "human in the loop solves it" rebuttal before anyone offers it.

Lee, H.-P., Sarkar, A., Tankelevitch, L. et al. (2025)

"The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers." CHI 2025. DOI: 10.1145/3706598.3713778

319 knowledge workers, 936 first-hand examples of GenAI use at work. Confidence in the AI predicted less critical thinking (B = -0.69); self-confidence in one's own expertise predicted more (B = +0.26). Effort shifts from information gathering to verification, from problem solving to response integration, from execution to what the authors call task stewardship.

Why it matters: moves the evidence from students to knowledge workers, which is FWC's actual client. "Task stewardship" is directly reusable language in organizational work.

Caveat: self-report. Cite it as a description of perceived effort reallocation, not of measured capability.

Bastani, H., Bastani, O., Sungu, A. et al. (2024)

"Generative AI Can Harm Learning." SSRN 4895486.

RCT, roughly 1,000 high school students. Unrestricted GPT-4 access raised practice performance ~48% and then lowered unaided exam performance ~17% relative to control. A safeguarded tutor version raised practice performance ~127% and eliminated the harm, while producing no measurable learning gain.

Why it matters: performance rose while capability fell, and students could not perceive the difference because during practice they were doing beautifully. The safeguarded arm is the proof that design determines the outcome, which is what keeps the paper from being a lament.

Cite the critique yourself: Tan, S. and Rajaratnam, V., "Critique of Generative AI Can Harm Learning Study Design," SSRN 4898213. You already handle Kosmyna this way, so the pattern is established and consistent.

Buçinca, Z., Malaya, M. B. and Gajos, K. Z. (2021)

"To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making." Proc. ACM Hum.-Comput. Interact. 5, CSCW1. arXiv:2102.09692

Cognitive forcing functions, meaning interface designs that require a person to commit to a judgment before the AI reveals its answer, reduce overreliance. They also reduce speed and satisfaction, and they work least well for the people with the lowest need for cognition, which is the equity problem sitting inside the intervention.

Why it matters: the existing, tested intervention. Its cost structure is itself a research question, and its differential effect by need-for-cognition is a direct bridge to Arielle's equity and neurodivergence work.

Fan, Y., Tang, L., Le, H. et al. (2025)

"Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance." British Journal of Educational Technology. DOI: 10.1111/bjet.13544

117 university students, four conditions: ChatGPT, human expert, writing analytics tool, control. ChatGPT users showed the largest essay score improvement, and no significant difference in knowledge gain or transfer. Same dissociation Bastani found, in a different population with a different task.

Why it matters: an independent replication of the core dissociation between visible performance and underlying capability. Two studies is a pattern; one is an anecdote with statistics.

Anthropic (Chen, Y., Benton, J. et al., 2025)

"Reasoning Models Don't Always Say What They Think." arXiv:2505.05410

Models were given hints that changed their answers, then examined for whether the chain of thought mentioned the hint. Claude 3.7 Sonnet verbalized it 25% of the time; DeepSeek R1, 39%. On the unauthorized-access hint type, 41% and 19% respectively. In reward-hacking environments, models exploited the shortcut more than 99% of the time while mentioning it in under 2% of reasoning traces.

Why it matters: the most human-readable interpretability artifact is the one most likely to be read as authentic when it is not. That last pair of numbers is the most quotable statistic available to this paper in any of its four versions. Use it in the talk.


2 of 4. Tier 2: The social science canonJenkins, H., Purushotma, R., Weigel, M.

Jenkins, H., Purushotma, R., Weigel, M. et al. (2006)

"Confronting the Challenges of Participatory Culture: Media Education for the 21st Century." MacArthur Foundation.

Source of the transparency problem, the participation gap, and the ethics challenge. The load-bearing empirical detail: students playing a history game absorbed the game's account of history as authentic because they "lacked a vocabulary to critique how the game itself constructed history."

Siemens, G. (2005)

"Connectivism: A Learning Theory for the Digital Age." International Journal of Instructional Technology and Distance Learning.

Learning resides in non-human appliances; competence is connection-making; "the pipe is more important than the content within the pipe." Steelman this rather than dismissing it. Its unaddressed gap is the paper's opening: connectivism has no account whatsoever of auditing the appliance.

Bainbridge, L. (1983)

"Ironies of Automation." Automatica 19(6), 775-779.

Automating a process erodes the operator's routine practice at precisely the moment the system depends on that operator for the cases automation cannot handle. Less practiced, more critical. Gives the argument an engineering pedigree rather than a humanities one, and Budzyń (2025) is now its measured confirmation in a new domain.

Eisner, E. (1979)

"The Null Curriculum." In The Educational Imagination.

What a school declines to teach still teaches, because the omission is a decision. Arielle's bridge to emergent misalignment. State it as a conjecture. No published result demonstrates that any feature direction corresponds to an omission in training data rather than to a presence.

Pike, K. (1967)

Emic and etic. Language in Relation to a Unified Theory of the Structure of Human Behavior.

The insider's categories versus the outside analyst's structural account. Arielle's strongest single contribution: probing classifiers are emic-adjacent and suffer the same limit as informant self-report, in that encoding a concept does not establish that the concept is used. Activation patching is etic, in that it intervenes and observes what changes rather than asking.

Clifford, J. and Marcus, G., eds. (1986)

Writing Culture: The Poetics and Politics of Ethnography.

The reflexive crisis. An ethnographic account can be coherent, compelling, publishable and still wrong: a narrative imposed on data rather than extracted from it. Structurally the same problem interpretability now calls faithful versus plausible. See also Geertz (1973) on thick description as an attempted discipline against exactly this.

Weak point to anticipate: Writing Culture concerned the ethnographer's authority over a described people who could object. Faithfulness in interpretability concerns correspondence to a computation. The ethical asymmetry does not transfer cleanly and a reviewer will say so. Address it in a footnote rather than hoping nobody notices.

boyd, d. (2014)

It's Complicated: The Social Lives of Networked Teens, and boyd on context collapse (drawing on Brandtzaeg and Lüders).

Information intended for one audience reaching unintended viewers, producing misinterpretation without any change to the content. Arielle's mapping to distributional shift in circuit interpretation is close to identity rather than analogy: a feature identified as meaning something specific in one prompt distribution does not reliably retain that meaning when the input distribution moves.

Nissenbaum, H., contextual integrity (via boyd)

Content is produced and interpreted in a context, and the context of creation may not be the context of consumption. Supplies the ethical register the paper otherwise lacks: faithful to whose context? Whoever decides which context counts as the real one when interpreting a feature is making a value judgment rather than a measurement. Interpretability does not currently ask this question at all.

Seaver, N. (2017)

"Algorithms as culture: Some tactics for the ethnography of algorithmic systems." Big Data & Society 4(2). DOI: 10.1177/2053951717738104

The essential citation for spine A's credibility. Seaver argues algorithmic systems should be treated as culture rather than as objects that culture surrounds, and offers concrete ethnographic tactics. Read this before drafting spine A, because if the anthropology-of-AI framing is already established territory, the novelty claim has to be narrowed to mechanistic interpretability specifically rather than AI generally.

Ito, M., Baumer, S., Bittanti, M., boyd, d. et al. (2010)

Hanging Out, Messing Around, and Geeking Out: Kids Living and Learning with New Media. MIT Press.

The connected-learning evidence base. In Arielle's Master's project it underwrites the informal-learning argument. For this paper it supplies the empirical account of what genuinely participatory technical learning looks like, which is the standard against which to judge whether open-sourced interpretability tooling (circuit-tracer, Neuronpedia, Gemma Scope) actually lowers the barrier or reproduces the participation gap one level up.

Ferdman, A. (2026)

"AI deskilling is a structural problem." AI & Society 41(4). DOI: 10.1007/s00146-025-02686-z

Introduces capacity-hostile environments: socio-technical conditions in which AI mediation disrupts the gradual habituation through which skill develops. Argues deskilling is a design property of environments rather than a failure of individual virtue, because we cannot expect people to be "virtuous superheroes."

Why it matters: the ready-made theoretical anchor for institutional legibility. It performs exactly the move the paper needs, which is relocating the unit of analysis from the user to the environment. It is also recent enough that citing it signals currency.

Aydeniz, S. (2026)

"Engineering Education at the Intersection of Knowledge, Work, and Power: Generative AI and Epistemic Authority." Sociology Compass. DOI: 10.1111/soc4.70230

Read this before committing to the term "interpretive authority." It is the nearest neighbour in the literature and it is in the education domain. Either it makes the construct redundant, or it is a citation that establishes the construct is live and needs the extension you are offering.

Akgün, S., Choi, K. and Lee, H. R. (2026)

"Designing a critical AI literacy program for K-8 STEM education: adopting a community-centered approach." Disciplinary and Interdisciplinary Science Education Research. DOI: 10.1186/s43031-026-00155-1

Participatory design with elementary students and public librarians, co-creating modules around community-relevant AI ethics. Three-part framework: asset-based views of learners, centering community knowledge, fostering critical consciousness.

Why it matters: the closest existing model for the Pittsfield community event, and a citable precedent for treating a public event as legitimate design research rather than outreach. Given Arielle's IRB certification, this is the template to copy.


3 of 4. Tier 3: Technical support and honest limitsTurpin, M., Michael, J., Perez, E.

Turpin, M., Michael, J., Perez, E. and Bowman, S. (2023)

"Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388

The original CoT faithfulness result. Pair with the Anthropic 2025 follow-up above.

NeurIPS 2025, "Revising and Falsifying Sparse Autoencoder Feature Explanations"

Automated SAE feature explanations are frequently too broad and blind to polysemanticity. The authors propose similarity-based falsification with hard negatives, a structured description format, and tree-based iterative refinement.

Why it matters: the honest limit on "point to the feature." Interpretability's own explanation artifacts do not currently survive falsification testing reliably. This is simultaneously the strongest support for the reflexive argument and the reason spine C cannot be the spine.

Kosmyna, N. et al. (2025)

"Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task." arXiv:2506.08872

Deliberately excluded from the evidence base. Cite it only alongside Stankovic et al. (2026), "Comment on: Your Brain on ChatGPT," arXiv:2601.00856, which challenges sample size, reproducibility, EEG methodology and reporting consistency.

Why it earns a paragraph anyway: it is the study everyone in the room will raise, and knowing its status makes you the most credible person there. It is also a live specimen of the phenomenon under study, in that a legible-sounding finding circulated far past what its evidence supported.


4 of 4. Still to acquire1.
  1. A systematic pass on the XAI overreliance literature. Finite and reviewable. Schemmer et al. on appropriate reliance, and the CHI 2025 partial-explanations work (DOI 10.1145/3710946) are the next two to pull.
  2. The classic automation bias literature. Parasuraman and Riley (1997) on use, misuse, disuse and abuse; Skitka, Mosier and Burdick (1999) on automation bias in cockpit tasks. These predate the AI framing entirely and give the historical section teeth it currently lacks.
  3. A citation-level check on spine A's central claim. Whether interpretability method papers already invoke ethnographic vocabulary. Forty to sixty papers, coded. Boring, finite, and it is the difference between a claim and an assertion.
  4. An existing validated scale to anchor the instrument against. Team psychological safety, or a decision-quality measure. Face validity alone will draw a methods objection.

Sections in this document