Working artifact for Palmer Foote and Arielle Pink

Fluency Without Understanding

A twenty year old media literacy problem, a forty year old automation problem, and the new science of AI interpretability are all describing the same failure. Mapping that overlap is a paper worth writing.

August 17, 2026 Source texts: Jenkins et al. (2006), Siemens (2005) For the interpretability block
The claim to put on the table

Henry Jenkins named the transparency problem in 2006. People become fluent users of a system while losing the ability to see how that system shapes what they perceive. Generative AI did not introduce this problem. It industrialized it.

The sharpest version of the argument is reflexive. The tools built to solve opacity (explanations, chain of thought, interpretability dashboards) have been empirically shown to deepen misplaced trust. Legibility and understanding come apart. That gap is the paper.

Figure 1

The connectivity map

Left column: what the social sciences already established. Right column: what interpretability research is finding now. Center: the concepts that carry across. Click any node for the detail and the citation.

Social science and media literacy Shared mechanism AI interpretability click a node

Select a node

Figure 2

The transparency problem has a long history

Your instinct to come at this historically is the right one, and it is also the strategic move. It converts "AI is making us dumb," a tired claim nobody trusts, into "here is a recurring structural pattern in tool adoption, and here is what changed this time." The second version is publishable.

c. 370 BCE

Plato, Phaedrus, the myth of Theuth

King Thamus rejects the gift of writing. It will "create forgetfulness in the learners' souls, because they will not use their memories." Writing gives students "the appearance of wisdom, not true wisdom." They will seem to know much while knowing nothing. This is the transparency problem in its first recorded form, and it is worth opening the paper with, because it disarms the reader. Yes, every generation says this. Our job is to say what is different now.

1983

Bainbridge, "Ironies of Automation"

The foundational automation safety paper. The irony: automating a process removes the operator's routine practice, so their skill degrades, precisely while the system increasingly depends on that operator to catch the cases automation cannot handle. Automation leaves the human less practiced and more critical at the same time. Directly transferable, and it gives the argument an engineering pedigree rather than a humanities one.

2005 to 2006

Siemens and Jenkins, the networked turn

Siemens: knowledge now lives in the network and in "non-human appliances," and competence becomes connection making. Jenkins: participatory culture is arriving faster than the literacies needed to navigate it, and he names the participation gap, the transparency problem, and the ethics challenge. Both describe the preconditions for what LLMs did next, twenty years early, without the vocabulary for it.

2021

The XAI reversal

Bansal et al. find that AI explanations increase the chance a person accepts the AI's recommendation regardless of whether it is correct. The field building transparency tools discovers that transparency tools can manufacture compliance. This is the hinge of the whole argument.

2024 to 2025

Measurement arrives

Bastani et al. put a number on skill erosion in a randomized trial. Vaccaro et al. show in meta-analysis that human and AI teams frequently underperform the better of the two alone. Lee et al. find that confidence in the AI predicts less critical thinking. The anecdote becomes an evidence base.

2026

Interpretability goes mainstream, without a theory of the user

Mechanistic interpretability matures into a serious science of the artifact, while chain of thought faithfulness remains an open question. What the field still lacks is any account of what happens to a person or an organization receiving an explanation. That absence is the opening.

Figure 3

Three levels of legibility, and which one is empty

The word "interpretability" is doing three different jobs. Separating them is the cleanest contribution the paper can make, because it shows the vacancy rather than arguing for it.

1

Model legibility: what is happening in the weights

Features, circuits, sparse autoencoders, attribution graphs. Object of study: the artifact. Audience: researchers. Well funded, concentrated in a handful of labs, and genuinely advancing.

Crowded. Not your fight.
2

Interaction legibility: can this person tell how the system shaped this output

The XAI and human factors tradition. Object of study: the human and system pair. Findings here are uncomfortable. Explanations frequently increase reliance without improving accuracy.

Active, and largely negative results
3

Institutional legibility: can a group see how the tool changed its own judgment

How a team's norms, standards of evidence, and distribution of expertise shift after adoption. Object of study: the collective. This is Jenkins and Siemens territory, it is where Form Wave Collective already works, and almost nobody is doing it empirically.

Nearly vacant. The opening.

The receipts

Five studies that carry the argument

Bring these. The difference between an interesting conversation and a paper is whether the claims have numbers attached.

StudyFindingWhy it matters here
Bastani et al.2024. RCT, roughly 1,000 students Unrestricted GPT-4 access raised practice scores +48%, then lowered unaided exam scores -17% against control. A safeguarded tutor version raised practice performance +127% and eliminated the harm, while producing no learning gain. The transparency problem with a number on it. Performance rose while capability fell, and the students could not perceive the difference, because during practice they were doing beautifully. This is the single most persuasive result you have.
Bansal et al.CHI 2021 AI explanations "increased the chance that humans will accept the AI's recommendation, regardless of its correctness." Explanations helped when the AI was right and hurt when it was wrong, for near zero net gain over showing a confidence score. The reflexive turn. Interpretability output is itself a persuasive medium. This is what licenses applying media literacy theory to interpretability tools rather than only to AI outputs.
Vaccaro, Almaatouq and MaloneNature Human Behaviour, 2024 Meta-analysis of 106 studies and 370 effect sizes. Human and AI combinations performed worse than the best of human or AI alone (g = -0.23). Losses concentrated in decision tasks (g = -0.27). Gains appeared in creation tasks (g = +0.19). Kills the easy "human in the loop solves it" rebuttal before anyone offers it. Also gives you a real design principle: the loop helps for making things and hurts for choosing between things.
Lee et al. (Microsoft Research and CMU)CHI 2025. 319 knowledge workers Confidence in GenAI predicted less critical thinking (B = -0.69). Self-confidence predicted more (B = +0.26). Effort shifted from information gathering to verification, from problem solving to response integration, from execution to "task stewardship." Moves the argument from students to knowledge workers, which is FWC's actual client. The stewardship framing is directly usable in org work.
Turpin et al., and Anthropic CoT faithfulness work2023, 2025 Chain of thought explanations can systematically misrepresent the actual reasons for a model's output. Models "don't always say what they think." The most human readable interpretability artifact is the one most likely to be read as authentic when it is not. Jenkins's exact structure, appearing in the technical literature.

Deliberately excluded: the MIT Media Lab "Your Brain on ChatGPT" EEG study. It is the one everyone will bring up, and a published commentary has challenged its sample size, reproducibility, EEG methodology, and reporting consistency. Knowing this makes you the most credible person in the room. It is also a perfect live example of the phenomenon: a legible sounding finding that circulated far past what its evidence supports.

How to run the conversation

Framing moves for the brainstorm

Open with the reframe, not the worry

  • Avoid opening with "AI is degrading critical thinking." Everyone has heard it, most people have already decided how they feel, and it sounds like moral panic.
  • Open with the naming. Jenkins gave this a name in 2006, and he was writing about video games and classrooms. The structure transfers exactly. That is a discovery, and discoveries are more interesting than complaints.
  • Then the reflexive hook. The tools built to fix opacity have been measured making trust worse. That is the sentence that makes people lean in.
"The transparency problem isn't new. What's new is that we finally have instruments precise enough to measure it, and the instruments have the same problem."

Give the vocabulary early

  • Separate the three legibilities (Figure 3) in the first ten minutes. Most confusion in these conversations comes from two people using "interpretability" to mean different levels.
  • Name what you are not claiming. You are not doing mechanistic interpretability and you are not competing with it. Saying so buys enormous credibility.
  • Use "fluency without understanding" as the recurring phrase. It is neutral, precise, and it avoids moralizing.
"We're not arguing the models are opaque. We're arguing the relationship is opaque, and nobody's measuring that one."

Keep the historical frame honest

  • Concede Plato immediately. Every technology has drawn this critique and most of those critiques were wrong. Say it before anyone else does.
  • Then earn the difference. Writing externalized memory. Search externalized retrieval. Generative AI externalizes the production of the reasoning itself, and returns it in fluent prose that carries no signal about its own reliability.
  • Bainbridge is the hinge. It is a case where the worry proved correct, was well documented, and led to real design changes in aviation and process control. The pattern is sometimes right, and it can be engineered around.

Land on construction, not critique

  • Point at the safeguarded tutor. In Bastani, changing the prompt design eliminated the harm entirely. Design determines whether offloading costs you the skill.
  • Cognitive forcing functions, meaning deliberately making a person commit to a judgment before the AI shows its answer, are an existing and testable intervention with a known cost in speed and satisfaction. That tradeoff is itself a research question.
  • Editability as the practical form of interpretability. The Intelligence Meridian thread. The offer becomes "here is my reasoning, change it."

Prepared

The four objections you will get

"This is just the Plato argument again. Every generation panics about a new technology."

Agreed, and that is why the paper opens there rather than hiding it. The claim is narrow and testable. Certain designs produce measurable capability loss that the user cannot perceive, and we now have randomized evidence of exactly that. Siemens argues persuasively that offloading is the condition of modern competence, so the argument is about which designs cost you the skill. Plato had an intuition. Bastani had a control group.

"Interpretability is a technical field. What do media studies have to offer it?"

Interpretability produces artifacts that get shown to people: dashboards, feature descriptions, reasoning traces. The moment an explanation has a human audience it becomes a representation, with conventions, omissions, and persuasive force. Media studies is the discipline that studies exactly that, and the CHI literature has already measured the effect. Explanations increase acceptance regardless of correctness. This is a documented failure the technical field cannot diagnose with its own tools.

"Isn't the answer just keeping a human in the loop?"

The meta-analytic evidence says no. Human and AI combinations underperformed the better of the two alone across 106 studies, with losses concentrated precisely in decision tasks, which is where "human in the loop" is usually invoked. The loop is a design surface with known failure modes, and treating it as a solution is what gets people into trouble. That finding is itself an argument for the paper.

"Aren't you just going to reproduce the doom literature?"

The tell is that we are excluding the most viral study in the space on methodological grounds, and that our historical section concedes the critique has usually been wrong. The paper's disposition is diagnostic. Here is a recurring structural pattern, here is what is genuinely different this time, here is which design choices change the outcome. Bastani's safeguarded arm is the proof that this is an engineering problem rather than a lament.

The pitch

What the collaborative paper could actually be

Shape of the thing

  • The Second Transparency ProblemPositions the paper as extending Jenkins rather than repeating him. Most defensible.
  • Legible to Whom? Interpretability Beyond the ModelForegrounds the three levels contribution. Strongest for an FAccT or CHI audience.
  • Fluency Without Understanding: What Media Literacy Research Knows That Interpretability Doesn'tMost provocative. Best for a talk or the community event, riskiest for peer review.
Core contribution
  1. Formalize the three levels of legibility as a framework, and show that levels 2 and 3 are underserved.
  2. Establish the historical lineage, from Phaedrus through Bainbridge and Jenkins to the 2021 to 2025 empirical turn, as one continuous problem rather than a series of moral panics.
  3. Argue the reflexive case. Interpretability artifacts are media, and they inherit media's failure modes.
  4. Propose institutional legibility as a measurable construct, with a first instrument.
What you would need to do
  • A systematic pass on the XAI overreliance literature. That is the empirical spine, and it is finite.
  • Decide between a pure position paper and a position paper plus a small original study. A short organizational instrument built from Jenkins's five conditions would make it publishable at a stronger venue.
  • Pick the venue early. CHI, CSCW, and FAccT reward exactly this kind of cross-disciplinary framing. A pure interpretability venue will not.
  • Divide labor honestly. Whoever owns the social science canon owns sections 1 and 2. Whoever owns the technical literature owns section 3.

For the brainstorm

Open questions worth putting on the whiteboard

  1. If explanations increase acceptance regardless of correctness, what is an explanation actually for, and who is it for?
  2. What would it take to measure institutional legibility? What is the survey item that detects a team whose standards of evidence have quietly moved?
  3. Is "editability" a genuinely different property from explainability, or a nicer interface on the same thing? What experiment would tell them apart?
  4. Bastani's safeguards worked. What is the general principle behind them, and does it survive contact with adults who can simply route around the safeguard?
  5. Vaccaro found losses in decision tasks and gains in creation tasks. Does the transparency problem apply to creative work at all, or is it specifically a pathology of judgment?
  6. Who is the paper for: the interpretability community, the learning sciences, or organizations trying to adopt this well? Each implies a different paper, and choosing is the first real decision.