Working artifact for Palmer Foote and Arielle Pink

Choosing the Spine

Three candidate papers sit in the material you and Arielle have already written. They look like variations on one idea. They are three different objects of study, and picking one is the decision that unblocks everything else.

August 18, 2026 Scope agreed: position paper plus a small instrument Prepared for the working session
What the three candidates have in common

Every version of this paper is asking who is authorized to say what a system means. Arielle asks whether the insider's account or the analyst's account governs. Your transparency problem asks whether the user can see the construction at all. The null curriculum asks whether an absence can be made to speak. One construct sits under all three: interpretive authority, and where it moves when a tool arrives.

That construct is also what the instrument should measure. If you build a survey that detects where interpretive authority relocated inside a team after AI adoption, it serves whichever spine you pick. So the instrument work can start before the spine decision is final, which means the decision does not have to be made today.

The candidates

Three papers, argued at full strength

Each tab makes the strongest available case for that spine, then the strongest case against it. The case against is the useful part. A spine you cannot defend under hostile questioning at CHI or FAccT will cost you a year.

A

Interpretability researchers have re-derived twentieth century social science

Arielle's frame. Owner of record: Pink
The claim

Mechanistic interpretability is fieldwork on an emergent system whose participants cannot report their own rules. Its methodological toolkit has independently reproduced distinctions that anthropology named decades ago: emic versus etic, faithful versus plausible interpretation, formal versus informal schema, and context collapse. The field is solving hard problems without the vocabulary that was already built for them.

Why this is the strongest candidate

  • Four independent bridges, not one analogy. Emic/etic maps onto probing versus activation patching. Faithfulness maps onto the Writing Culture crisis. Formal/informal schema maps onto interpretable-by-design versus superposed representations. Context collapse maps onto distributional shift. Any one of them could be coincidence. Four is a pattern.
  • The emic/etic mapping is technically precise. Probing asks whether a concept is encoded, which is close to self-report, and a probe can succeed on a feature the model never uses. Activation patching intervenes and observes what changes. That is exactly the difference between what an informant says the rule is and what behavior reveals the rule to be.
  • Hard to scoop. The people with the technical literature rarely have Pike, Geertz and Clifford. The people with the anthropology rarely have transcoders and attribution graphs. That intersection is thin, and thin intersections stay open for a while.
  • It arrives with a gift, not a complaint. Anthropology spent forty years learning that a coherent, publishable, compelling ethnographic account can still be wrong. Interpretability is discovering that same trap right now under the name faithfulness. The transferable methodological caution is a real offer to the field.

Where a reviewer will push

  • The central claim is sociological and currently unevidenced. "Researchers independently re-derived X" is an assertion about what a community says and cites. Right now you have four apt readings. A skeptic will ask for the corpus.
  • Tone risk. A paper telling a technical field it has been doing amateur anthropology reads badly if the framing slips even slightly. It has to arrive as a loan of vocabulary, and the abstract has to establish that in the first three sentences.
  • The faithfulness bridge is the weakest of the four. Writing Culture was about the ethnographer's authorial power over a described people who could object. Interpretability's faithfulness problem is about correspondence to a computation. The ethical asymmetry does not transfer cleanly, and someone will say so.
  • Palmer carries technical risk here. If the probing versus patching account has any imprecision, the whole paper is discredited by the one paragraph a technical reviewer reads closely.
Unit of analysis
The research field, as a culture
Venue fit
FAccT, Big Data & Society, Science Technology & Human Values
Missing work
Coded read of 40 to 60 interpretability method papers
Time to draft
Longest. Three to five months
What would kill it

An interpretability researcher publishing a methods paper that explicitly cites Pike or Geertz. Check first. If the vocabulary is already arriving in the field, the claim inverts from discovery to description and the paper loses its edge overnight.

B

Fluency without understanding, and the three levels of legibility

Palmer's frame. Owner of record: Foote
The claim

Jenkins named the transparency problem in 2006: people become fluent users of a system while losing the ability to see how it shapes what they perceive. The word interpretability is now doing three separate jobs, at the level of the model, the interaction, and the institution. Funding sits almost entirely at the first. Human consequence sits almost entirely at the second and third.

Why this is the safest candidate

  • The evidence base already exists and it is finite. Bansal, Buçinca, Vasconcelos, Vaccaro, Lee, Bastani, Budzyń. You could have a defensible systematic pass done in six weeks.
  • Budzyń et al. (2025) changed the argument's weight class. Nineteen experienced endoscopists, adenoma detection rate falling from 28.4 percent to 22.4 percent on unassisted procedures after routine AI exposure. Bainbridge's 1983 irony, measured in live clinical practice, in a Lancet title. That is no longer a humanities argument.
  • The three-level taxonomy is a clean, citable contribution that solves a real problem: most confusion in these rooms comes from two people using one word for three different objects.
  • It is directly monetizable. Level three is FWC's existing consulting surface. The instrument becomes a client diagnostic the week the paper is drafted.

Where a reviewer will push

  • The reflexive hook is no longer clean. Vasconcelos et al. (2023), five studies and 731 participants, showed explanations do reduce overreliance when the cost of engaging with the explanation drops below the cost of doing the task yourself. "Explanations manufacture compliance" is now a special case, not a finding. You have to update the hook or a reviewer will update it for you.
  • XAI overreliance is a crowded room. Everything up to the three-level taxonomy has been said. The taxonomy is the contribution, and taxonomies without data get desk-rejected at CHI.
  • The historical lineage section is the part most likely to read as familiar. Plato through Bainbridge through Jenkins is a good talk and a thin paper section.
  • Bastani has a published design critique. Tan and Rajaratnam went after it on SSRN. Cite it yourself before someone else does.
Unit of analysis
The person and the organization
Venue fit
CHI, CSCW, FAccT, Learning Media and Technology
Missing work
The instrument. Without it this is a framework note
Time to draft
Shortest. Six to ten weeks
What would kill it

Nothing kills it. Its risk is the opposite one: it gets accepted somewhere respectable and changes no conversation, because the taxonomy is useful and the argument is familiar. Ask yourself whether "make a huge blow in the scene" and "safest candidate" can be the same sentence.

C

A null curriculum you can point to

Arielle's boldest move. Shared ownership
The claim

Eisner's null curriculum holds that what a school declines to teach still teaches, because the omission is itself a decision. Emergent misalignment is the mechanistic version: fine-tuning on insecure code produced unrelated toxic persona features that nobody trained for. If persona directions are locatable in activation space, then for the first time in the history of either field, a null curriculum might be a thing you can point at inside a system.

Why this is the one people will repeat

  • It is the only idea here that is stage-ready. Concrete, slightly unsettling, graspable by a non-technical audience in one sentence. That is the Pittsfield event, the TEDx talk, and the Berkshire Muse room.
  • It converts a curriculum-theory concept into an empirical question, which is exactly the kind of move that gets a concept adopted outside its home discipline.
  • It gives the education framing something the other two lack: a reason for interpretability researchers to care what curriculum theorists think, rather than the other way around.

Where a reviewer will push, hard

  • The empirical claim is not yet true, and the gap is findable in thirty seconds. Emergent misalignment shows a persona feature forming from what the model was shown. The null curriculum is about what it was not shown. Nobody has demonstrated that any feature direction corresponds to an omission. Claiming otherwise is the single fastest way to lose a technical audience.
  • It is one idea, and one idea is a section. There is not enough here to carry twelve pages without either A or B underneath it.
  • The falsification problem is live. A NeurIPS 2025 result found that automated sparse-autoencoder feature explanations are frequently too broad and blind to polysemanticity, which means "point to the feature" is doing more work than the tooling currently supports.
Unit of analysis
The model
Venue fit
The talk, the event, the popular essay
Missing work
An actual experiment, or an honest downgrade to a conjecture
Time to draft
As a section: two weeks. As a paper: do not
Recommendation on C

Do not make this the spine. Make it the closing section of whichever spine you pick, framed explicitly as a conjecture with a stated experiment attached. Stated as a conjecture it is the most memorable thing in the paper. Stated as a finding it is the thing that gets the paper rejected.

D

Interpretive authority: where the right to say what a system means relocates

The construct underneath all three. Genuinely co-authored
The claim

Adopting an AI system moves interpretive authority. It moves from the practitioner to the tool, from the team's shared standard of evidence to whatever the interface presents as sufficient, and from the people affected by a judgment to the people who build the explanation of it. That movement is measurable, it happens at three scales, and no existing literature measures it at the scale that matters most.

Why this is more than a compromise

  • It explains why the three candidates feel like one idea. Emic versus etic is a fight over whose account governs. The transparency problem is a loss of standing to question the account. The null curriculum asks whether an absence has standing to be read at all. Same construct, three scales.
  • It makes the instrument central rather than bolted on. The survey measures a named construct instead of a vague sense that a team changed. That is the difference between a framework paper and an empirical one.
  • It gives Arielle and Palmer non-overlapping halves that both matter. Arielle owns the theory of authority, from Pike and Geertz to Nissenbaum's contextual integrity and boyd's context collapse. Palmer owns the measurement and the technical account. Neither half is decorative.
  • Ferdman's "capacity-hostile environments" (AI & Society, 2026) is a ready-made theoretical anchor. She argues deskilling is structural rather than a failure of individual virtue, which is the same move you need: the unit of analysis is the environment, not the user.
  • Contextual integrity supplies the ethical register. Faithful to whose context? Whoever decides which context counts as the real one when interpreting a feature is making a value judgment. That question is not currently asked in the interpretability literature at all.

Where a reviewer will push

  • Synthesis papers can read as evasive. If the construct is not doing real work by page three, it looks like three papers stapled together with a new noun on the cover.
  • You have to operationalize it or it is just a better metaphor. The whole bet is the instrument. If the survey does not discriminate between teams, the paper collapses back to B.
  • It needs a name that is not already taken. Check epistemic authority, epistemic dependence, and interpretive labor in philosophy and STS before committing. Aydeniz (Sociology Compass, 2026) is already using "epistemic authority" for generative AI in engineering education.
Unit of analysis
All three, held together by one construct
Venue fit
FAccT or CSCW. Both reward this shape
Missing work
The instrument, plus a terminology check
Time to draft
Ten to fourteen weeks
Honest caveat

This is the option a thought partner proposes and an author has to decide is real. It is not a way of avoiding the choice. If the construct does not feel alive to both of you within one conversation, pick B and ship it.

Side by side

The four criteria that should actually decide this

A. AnthropologyB. TransparencyC. Null curriculumD. Interpretive authority
Novelty High. Thin intersection Moderate. Crowded room High, and unproven High if the construct holds
Defensibility under hostile questioning Medium. Needs a corpus High. Evidence exists Low. Known gap Medium to high
Reviewers who exist for it FAccT, STS journals CHI, CSCW, FAccT Few. It is a talk FAccT, CSCW
Survives a big lab publishing nearby Yes. Labs do not write this Partly. Overreliance work is active No. One paper ends it Yes
Division of labor Arielle carries 70 percent Palmer carries 70 percent Uneven and thin Roughly even
Doubles as FWC client work Weakly Directly As a talk Directly

Before the next draft

Four things in the current map need fixing

These came out of checking every number in transparency_problem_map_1.html against primary sources. Three are corrections. One is an upgrade.

The reflexive hook is overstated, and fixing it makes the argument better

The map presents Bansal et al. (2021) as showing that explanations manufacture compliance, full stop. Vasconcelos et al. (CSCW 2023) ran five studies with 731 participants and found explanations do reduce overreliance, specifically when engaging with the explanation costs less effort than doing the task yourself. Their reading of the earlier null results is that the explanations simply failed to lower verification cost. Why this helps you: the claim upgrades from "explanations are bad" to "an explanation only works when reading it is cheaper than checking the work." That is a design principle, an education principle, and a testable prediction. It is also a much harder sentence for a reviewer to knock down.

Vaccaro's creation-task gain is not statistically significant

The map reports "Gains appeared in creation tasks (g = +0.19)." The published confidence interval is -0.09 to 0.48, which crosses zero. The overall effect (g = -0.23, CI -0.39 to -0.07) and the decision-task effect (g = -0.27, CI -0.44 to -0.10) are both solid. The creation-task result is directional at best. Stating it as a finding is exactly the kind of thing that costs credibility in a room where someone has read the paper. Corpus: 370 effect sizes from 106 experiments.

Bastani needs a defensive citation

Tan and Rajaratnam published a design critique of "Generative AI Can Harm Learning" (SSRN 4898213). The study is still your most persuasive single result, and citing the critique yourself is the move that protects it. You already do this correctly with the Kosmyna EEG study, so the pattern is established.

Two studies published since the map should go in immediately

Budzyń et al., Lancet Gastroenterology & Hepatology (2025). Nineteen experienced endoscopists across four Polish centres, each with more than 2,000 prior procedures. Adenoma detection rate on non-AI-assisted colonoscopies fell from 28.4 percent (226/795) before routine AI exposure to 22.4 percent (145/648) after. A 20 percent relative decline in the unaided skill of experts. This is Bainbridge measured in live clinical practice, and it is the strongest evidence in the whole field that the pattern is real outside a classroom.

Anthropic (2025), "Reasoning Models Don't Always Say What They Think." Claude 3.7 Sonnet verbalized the hint it actually used 25 percent of the time; DeepSeek R1, 39 percent. In the reward-hacking condition, models exploited the shortcut more than 99 percent of the time and mentioned doing so in under 2 percent of their reasoning traces. That last pair of numbers is the most quotable statistic available to this paper, in any of its four versions.

The instrument

What a survey of institutional legibility has to do

You chose position paper plus a small instrument. The instrument is what moves this from a framework note to something a venue takes seriously, and it is also the only part that works as billable FWC diagnostic. A first draft lives in the 02_Evidence folder. Its design problem is stated here.

What it has to detect

  • A moved standard of evidence. The team now accepts as sufficient something it would have questioned eighteen months ago, and nobody voted on that.
  • A redistributed expertise map. Who gets asked has changed. The person who used to be consulted is now consulted after the tool, or not at all.
  • Loss of the unaided baseline. Nobody in the group could now say how good the group is without the tool, because nobody has measured it since adoption. Budzyń is the reason this item exists.
  • Explanation as terminal. How often does a generated rationale end a discussion rather than start one.

The hard design problems

  • Self-report is the wrong instrument for an unperceived change. The entire construct is that people cannot see it. Items have to ask about observable practice, not perception. "When did this team last review a decision the tool made?" beats "Do you rely on AI too much?"
  • You need a retrospective anchor. Some version of before and after, which is fragile. A behavioral proxy would be stronger if you can find one.
  • Arielle is IRB certified. That is not a small asset. It makes real deployment at the community event legitimately feasible, and it is worth deciding early whether you want that data.
  • Validate against something. Even a small correlation with an existing measure of team psychological safety or decision quality would let you claim construct validity instead of face validity.

If you want a recommendation

What I would do, and why you might not

Draft toward D and hold B as the fallback. Spend one conversation, ninety minutes, testing whether interpretive authority feels like a real construct to both of you or like a clever noun. If it holds, you have a genuinely co-authored paper where neither half is decoration, and the instrument becomes the centerpiece instead of an appendix. If it does not hold by the end of that conversation, write B, put Arielle's emic/etic section in as the theoretical frame, close with the null curriculum as an explicit conjecture, and get it out in ten weeks.

The reason to overrule me: A is the only one of the four that no other pair of people on earth is currently positioned to write. If "make a blow in the scene" is the actual objective rather than a first publication, A is the higher-variance, higher-ceiling bet, and the corpus work it needs is finite and boring rather than uncertain.

  • Interpretive Authority: What Moves When a Team Adopts an Interpretable SystemFor D. Foregrounds the construct and the instrument. Best FAccT or CSCW fit.
  • Legible to Whom? Interpretability Beyond the ModelFor B or D. Carries the three-level contribution. Broadest reach.
  • Fieldwork on a Grown System: Interpretability's Unnamed AnthropologyFor A. Establishes the loan-of-vocabulary framing in the title itself, which defuses the tone risk.
  • The Null Curriculum of a Language ModelFor the talk and the Pittsfield event. Do not use it on the paper.
Next three things, in order
  1. Ninety-minute conversation with Arielle testing construct D. Nothing else gets decided first.
  2. Terminology check on "interpretive authority" against epistemic authority, epistemic dependence, and interpretive labor. Aydeniz (2026) is the nearest neighbour.
  3. Draft ten instrument items. Ten, not thirty. Items that ask about observable practice.
Decisions you can defer
  • Venue. All four candidates keep FAccT and CSCW open.
  • Whether the community event supplies data. That depends on the instrument existing first.
  • The historical lineage section. It is written in your head already and it fits any spine.
  • The title. Titles are the last decision, not the first one.

For the working session

Questions that actually resolve the choice

  1. Is the paper's object the field, the user, the organization, or the model? Answer this and three of the four candidates disappear on their own.
  2. Does "interpretive authority" survive being stated back to you in one sentence by the other person? If either of you has to reach for a second sentence, the construct is not ready.
  3. What is the first thing you would say to an interpretability researcher who asks what social science offers them? Whichever spine gives you a fast, non-defensive answer to this is the spine that will survive a conference Q and A.
  4. If a lab published something adjacent in six months, which candidate would still be worth writing? This is the question that separates the high-ceiling bet from the safe one.
  5. Which of you wants to own the instrument? Whoever owns it owns the paper's empirical claim, and that shapes author order and the technical burden.
  6. Is the target one paper or a paper plus a talk plus a community event with different arguments in each? The honest answer is probably the third, and saying it out loud stops you from cramming the vivid material into the peer-reviewed version where it will get cut.