A twenty year old media literacy problem, a forty year old automation problem, and the new science of AI interpretability are all describing the same failure. Mapping that overlap is a paper worth writing.
Henry Jenkins named the transparency problem in 2006. People become fluent users of a system while losing the ability to see how that system shapes what they perceive. Generative AI did not introduce this problem. It industrialized it.
The sharpest version of the argument is reflexive. The tools built to solve opacity (explanations, chain of thought, interpretability dashboards) have been empirically shown to deepen misplaced trust. Legibility and understanding come apart. That gap is the paper.
The connectivity map
Left column: what the social sciences already established. Right column: what interpretability research is finding now. Center: the concepts that carry across. Click any node for the detail and the citation.
Select a node
The transparency problem has a long history
Your instinct to come at this historically is the right one, and it is also the strategic move. It converts "AI is making us dumb," a tired claim nobody trusts, into "here is a recurring structural pattern in tool adoption, and here is what changed this time." The second version is publishable.
King Thamus rejects the gift of writing. It will "create forgetfulness in the learners' souls, because they will not use their memories." Writing gives students "the appearance of wisdom, not true wisdom." They will seem to know much while knowing nothing. This is the transparency problem in its first recorded form, and it is worth opening the paper with, because it disarms the reader. Yes, every generation says this. Our job is to say what is different now.
The foundational automation safety paper. The irony: automating a process removes the operator's routine practice, so their skill degrades, precisely while the system increasingly depends on that operator to catch the cases automation cannot handle. Automation leaves the human less practiced and more critical at the same time. Directly transferable, and it gives the argument an engineering pedigree rather than a humanities one.
Siemens: knowledge now lives in the network and in "non-human appliances," and competence becomes connection making. Jenkins: participatory culture is arriving faster than the literacies needed to navigate it, and he names the participation gap, the transparency problem, and the ethics challenge. Both describe the preconditions for what LLMs did next, twenty years early, without the vocabulary for it.
Bansal et al. find that AI explanations increase the chance a person accepts the AI's recommendation regardless of whether it is correct. The field building transparency tools discovers that transparency tools can manufacture compliance. This is the hinge of the whole argument.
Bastani et al. put a number on skill erosion in a randomized trial. Vaccaro et al. show in meta-analysis that human and AI teams frequently underperform the better of the two alone. Lee et al. find that confidence in the AI predicts less critical thinking. The anecdote becomes an evidence base.
Mechanistic interpretability matures into a serious science of the artifact, while chain of thought faithfulness remains an open question. What the field still lacks is any account of what happens to a person or an organization receiving an explanation. That absence is the opening.
Three levels of legibility, and which one is empty
The word "interpretability" is doing three different jobs. Separating them is the cleanest contribution the paper can make, because it shows the vacancy rather than arguing for it.
Features, circuits, sparse autoencoders, attribution graphs. Object of study: the artifact. Audience: researchers. Well funded, concentrated in a handful of labs, and genuinely advancing.
Crowded. Not your fight.The XAI and human factors tradition. Object of study: the human and system pair. Findings here are uncomfortable. Explanations frequently increase reliance without improving accuracy.
Active, and largely negative resultsHow a team's norms, standards of evidence, and distribution of expertise shift after adoption. Object of study: the collective. This is Jenkins and Siemens territory, it is where Form Wave Collective already works, and almost nobody is doing it empirically.
Nearly vacant. The opening.Five studies that carry the argument
Bring these. The difference between an interesting conversation and a paper is whether the claims have numbers attached.
| Study | Finding | Why it matters here |
|---|---|---|
| Bastani et al.2024. RCT, roughly 1,000 students | Unrestricted GPT-4 access raised practice scores +48%, then lowered unaided exam scores -17% against control. A safeguarded tutor version raised practice performance +127% and eliminated the harm, while producing no learning gain. | The transparency problem with a number on it. Performance rose while capability fell, and the students could not perceive the difference, because during practice they were doing beautifully. This is the single most persuasive result you have. |
| Bansal et al.CHI 2021 | AI explanations "increased the chance that humans will accept the AI's recommendation, regardless of its correctness." Explanations helped when the AI was right and hurt when it was wrong, for near zero net gain over showing a confidence score. | The reflexive turn. Interpretability output is itself a persuasive medium. This is what licenses applying media literacy theory to interpretability tools rather than only to AI outputs. |
| Vaccaro, Almaatouq and MaloneNature Human Behaviour, 2024 | Meta-analysis of 106 studies and 370 effect sizes. Human and AI combinations performed worse than the best of human or AI alone (g = -0.23). Losses concentrated in decision tasks (g = -0.27). Gains appeared in creation tasks (g = +0.19). | Kills the easy "human in the loop solves it" rebuttal before anyone offers it. Also gives you a real design principle: the loop helps for making things and hurts for choosing between things. |
| Lee et al. (Microsoft Research and CMU)CHI 2025. 319 knowledge workers | Confidence in GenAI predicted less critical thinking (B = -0.69). Self-confidence predicted more (B = +0.26). Effort shifted from information gathering to verification, from problem solving to response integration, from execution to "task stewardship." | Moves the argument from students to knowledge workers, which is FWC's actual client. The stewardship framing is directly usable in org work. |
| Turpin et al., and Anthropic CoT faithfulness work2023, 2025 | Chain of thought explanations can systematically misrepresent the actual reasons for a model's output. Models "don't always say what they think." | The most human readable interpretability artifact is the one most likely to be read as authentic when it is not. Jenkins's exact structure, appearing in the technical literature. |
Deliberately excluded: the MIT Media Lab "Your Brain on ChatGPT" EEG study. It is the one everyone will bring up, and a published commentary has challenged its sample size, reproducibility, EEG methodology, and reporting consistency. Knowing this makes you the most credible person in the room. It is also a perfect live example of the phenomenon: a legible sounding finding that circulated far past what its evidence supports.
Framing moves for the brainstorm
The four objections you will get
Agreed, and that is why the paper opens there rather than hiding it. The claim is narrow and testable. Certain designs produce measurable capability loss that the user cannot perceive, and we now have randomized evidence of exactly that. Siemens argues persuasively that offloading is the condition of modern competence, so the argument is about which designs cost you the skill. Plato had an intuition. Bastani had a control group.
Interpretability produces artifacts that get shown to people: dashboards, feature descriptions, reasoning traces. The moment an explanation has a human audience it becomes a representation, with conventions, omissions, and persuasive force. Media studies is the discipline that studies exactly that, and the CHI literature has already measured the effect. Explanations increase acceptance regardless of correctness. This is a documented failure the technical field cannot diagnose with its own tools.
The meta-analytic evidence says no. Human and AI combinations underperformed the better of the two alone across 106 studies, with losses concentrated precisely in decision tasks, which is where "human in the loop" is usually invoked. The loop is a design surface with known failure modes, and treating it as a solution is what gets people into trouble. That finding is itself an argument for the paper.
The tell is that we are excluding the most viral study in the space on methodological grounds, and that our historical section concedes the critique has usually been wrong. The paper's disposition is diagnostic. Here is a recurring structural pattern, here is what is genuinely different this time, here is which design choices change the outcome. Bastani's safeguarded arm is the proof that this is an engineering problem rather than a lament.
What the collaborative paper could actually be
Open questions worth putting on the whiteboard