v1.0  /  August 19, 2026

Institutional Interpretability Prompt Pack

A research query kit for Perplexity, Elicit, Consensus and Semantic Scholar

The construct under investigation

Institutional interpretability: the degree to which a team can perceive how a tool has changed its own norms, its standards of evidence, and its distribution of expertise.

Why these questions

Three levels, one word

The word interpretability is doing three jobs. Model interpretability asks what is happening in the weights. Interaction interpretability asks whether a user can tell how the system shaped a given output. Institutional interpretability asks whether a group can see how the tool changed its norms, its judgment, and who it asks. Funding sits almost entirely at the first level. Human consequence sits at the second and third. The third is close to vacant, and it is the one this article is for.

Perception of change, not quality of outcome

A team can have adapted well and still score low. That is the construct working correctly. Keep the two separate in every question, or the research measures satisfaction instead.

What this pack is really asking

Three questions sit under the twenty prompts. Is the territory open (Track 0). Does the phenomenon exist outside your own reading of it (Tracks 2 and 4). Can it be measured (Track 3). Track 5 tries to break whatever the others find, and Track 6 decides where it goes.

Contents

How to run these

Use deep research mode for Tracks 0 through 2.

Perplexity Deep Research, Elicit, or Consensus. The prompts are written long on purpose. Shallow search mode collapses them into a blog-post summary and drops the methodological detail that makes a source usable.

One prompt per thread.

Chaining unrelated prompts in a single thread carries framing forward and the engine starts agreeing with you. Open a new thread each time.

An empty result is a result.

Every prompt asks the engine to say plainly when it finds nothing. On Track 0 a genuine null is the most valuable outcome in the pack, because it means the territory is open.

Nothing enters the Evidence Ledger unverified.

These engines get DOIs, sample sizes and confidence intervals wrong at a rate that will embarrass you in review. P17 is the verification prompt. Run it on every figure before the figure goes anywhere near a draft.

Ask for the disconfirming version of anything that lands.

Track 5 is not optional cleanup. Run P16 against your own thesis once a month while drafting.

Label the source type yourself.

Peer-reviewed, preprint, standards document, consultancy white paper, vendor blog. The engines blur these constantly and the blur is where credibility gets lost.

The reusable scaffold

Every prompt in this pack is built on this shape. Use it when you need a query the pack does not cover. The last line is the one that does the most work.

Template
ROLE: You are a research librarian with a background in [FIELD]. Prioritize precision over coverage.

QUESTION: [THE ONE THING YOU WANT TO KNOW, IN A SINGLE SENTENCE]

SCOPE: Peer-reviewed literature and working papers, 2015 to 2026, plus foundational work of any date where it is directly load-bearing. Include arXiv and SSRN preprints and label them as such.

FOR EACH SOURCE RETURN: author, year, title, venue, DOI or stable link, method, sample size and composition, the single finding most relevant to the question with its actual numbers, and one sentence on its limits.

FORMAT: A table first. Then one short paragraph on what this body of work does not cover.

CONSTRAINTS: Distinguish peer-reviewed work from preprints and from vendor or consultancy output. Do not substitute adjacent literature for the thing I asked about. If the work I am describing does not exist, say so plainly and then name the nearest adjacent literature, labelled as adjacent.
0
Track 0

Is the territory open?

Run these three before anything else. Each one can end the article in its current form, which is exactly why they go first. A hit here saves you three months. Budget one afternoon.

P01RUN FIRST

Prior art on the term itself

Establishes whether institutional interpretability is a name you are coining or a name you are borrowing without knowing it.

Paste this
Search for existing scholarly and practitioner use of the term "institutional interpretability" and its near neighbours in the context of artificial intelligence and algorithmic systems: "organizational interpretability", "institutional legibility", "organizational legibility", "organizational explainability", "collective interpretability", "organisational transparency".

List every source that uses any of these as a named construct rather than in passing. For each: author, year, venue, DOI or stable link, the definition given, and whether the term is load-bearing in the argument or incidental.

Cover four categories and label which each source belongs to: peer-reviewed literature, arXiv and SSRN preprints, policy and standards documents (NIST, ISO, OECD, EU AI Act commentary), and consultancy or vendor frameworks.

If none of these terms is established anywhere as a named construct, say so directly rather than filling the answer with adjacent material. Then list the five closest existing constructs that occupy the same conceptual space, with citations.
Follow up in the same thread
  • Forward-cite the strongest hit. Who has used it since, and for what?
  • Search the same terms in French, German and Spanish scholarship. The construct may exist under a translation.
P02RUN FIRST

Is the anthropology framing already taken?

Spine A in the decision brief depends on this being open. If interpretability researchers are already citing Geertz and Pike, the claim inverts from discovery to description.

Paste this
I want to know whether the framing "mechanistic interpretability is a form of ethnography or cultural analysis" has already been published.

Search three things separately and report them separately:

(a) Interpretability or explainable-AI methods papers that cite anthropological or ethnographic theory by name. Look specifically for Geertz, Clifford and Marcus, Pike, "thick description", "Writing Culture", "emic", "etic", "informant".

(b) STS and anthropology-of-algorithms work that takes model internals as its object rather than the sociotechnical surroundings of a system. Start from Nick Seaver (2017), "Algorithms as culture: Some tactics for the ethnography of algorithmic systems", Big Data and Society, DOI 10.1177/2053951717738104, and forward-cite it.

(c) Any paper that uses emic and etic to describe the difference between probing classifiers and activation patching or causal intervention.

For each source: author, year, venue, DOI, and one sentence stating exactly what claim is made.

Finish with a direct verdict in one paragraph: is this framing open territory, partially occupied, or already established, and what is the single strongest piece of evidence for that verdict.
Follow up in the same thread
  • Same query restricted to the last 12 months. This field moves fast enough that a six-month-old null is stale.
  • Check the reverse direction: has any anthropologist published on transformer internals specifically?
P03RUN FIRST

Terminology check on interpretive authority

Named in the decision brief as blocking construct D. Aydeniz 2026 is already using the adjacent term in the education domain.

Paste this
Compare the following constructs as they are used in philosophy, sociology, science and technology studies, and education research. Tell me whether any of them already covers this idea: the right to say what a system or a situation means, and where that right sits inside an organization.

Constructs to compare: interpretive authority, epistemic authority, epistemic dependence, epistemic injustice (Fricker), interpretive labour (Graeber), cognitive authority (Wilson), and sensemaking authority.

For each: originating author and work, the canonical definition, the field where it is primarily used, and how it has been applied to AI or automation if at all.

Include and address directly: Aydeniz, S. (2026), "Engineering Education at the Intersection of Knowledge, Work, and Power: Generative AI and Epistemic Authority", Sociology Compass, DOI 10.1111/soc4.70230.

Finish with a direct verdict: is "interpretive authority" available as a new named construct, a redundant relabel of an existing one, or a defensible extension of one, and if an extension, of which and in what specific respect.
Follow up in the same thread
  • If it is a relabel, ask: what would have to be added to make it a genuine extension?
  • Ask for the three terms most likely to be objected to by a FAccT reviewer, and why.
1
Track 1

Naming the construct

Organizational theory has been circling this for fifty years without the AI case. Find the nearest existing construct before you defend a new one. The strongest position is an extension of something established, not a coinage.

P04CORE

Adjacent constructs in organizational theory

Identifies the competitor construct a reviewer will name in the first paragraph of their review.

Paste this
I am developing a construct called institutional interpretability, defined as: the degree to which a team or organization can perceive how a tool has changed its own norms, its standards of evidence, and its distribution of expertise. The emphasis is on perception of change rather than quality of outcome. A team can have adapted well and still be unable to see what changed.

Map this construct against established constructs in organizational theory. For each, tell me whether it is the same thing, an overlapping thing, or a different thing, and say why in one sentence:

- Organizational sensemaking (Weick)
- Absorptive capacity (Cohen and Levinthal)
- Organizational mindfulness and high-reliability organizing (Weick and Sutcliffe)
- Organizational routines, ostensive and performative aspects (Feldman and Pentland)
- Technological frames (Orlikowski and Gash)
- Imbrication (Leonardi)
- Organizational forgetting and knowledge depreciation (Argote)
- Normalization of deviance (Vaughan)
- Organizational reflexivity and team reflexivity (West)

For each: canonical citation with DOI, the definition, whether a validated measure exists and its citation, and the verdict.

End by naming the single strongest competitor to my construct, and state what my construct would have to do that this one does not do in order to earn its own name.
Follow up in the same thread
  • Take the strongest competitor and ask for every published critique of it.
  • Ask which of these constructs has a validated short-form scale, since that is the anchoring candidate.
P05CORE

The legibility lineage, inverted

Scott is already in the library. The move worth owning is the inversion: the institution becoming illegible to itself.

Paste this
Trace how James C. Scott's concept of legibility, from "Seeing Like a State" (1998), has been applied downward from the state to the organization and to the individual.

I want three things:

1. Work that uses legibility as an analytic term for what an institution can and cannot see about itself, as opposed to what it makes visible about its subjects.
2. Work applying Scott's legibility to algorithmic, data or AI systems.
3. Published critiques of the concept and its limits, including the charge that it is too capacious to be falsifiable.

For each source: author, year, venue, DOI, the specific move made with the concept, and the direction of the gaze, meaning whether the institution is looking at its subjects or at itself.

I am specifically interested in any work that inverts Scott: cases where the apparatus of measurement makes the institution less legible to itself rather than making others more legible to it. If no such work exists, say so plainly, because that inversion is the contribution I am considering.
Follow up in the same thread
  • Ask what Scott himself wrote about firms and organizations rather than states.
  • Cross-check against the 'metis' side of the argument: who has applied tacit knowledge loss to AI adoption?
2
Track 2

Does the phenomenon exist?

The evidence base already in the ledger is strong on individuals and thin on collectives. These four prompts hunt for the organization-level evidence. The distinction that matters throughout: a study that measured a group's capability directly, rather than summing individual scores.

P06CORE

Deskilling with the organization as the unit of analysis

The ledger has Budzyn, Bainbridge and Ferdman. This finds what is missing, and specifically anything measured collectively.

Paste this
Find empirical studies of skill loss, capability erosion or expertise degradation caused by automation or AI where the unit of analysis is the team, the organization or the profession rather than the individual worker.

I already have these and do not need them summarized: Budzyn et al. (2025) in The Lancet Gastroenterology and Hepatology on endoscopist deskilling; Bainbridge (1983) "Ironies of Automation"; Ferdman (2026) in AI and Society on capacity-hostile environments.

Find what I do not have. Search across medicine, aviation, maritime, law, finance, engineering, translation, journalism and software engineering.

For each study: author, year, venue, DOI, design, sample size and composition, the effect with its actual numbers and confidence interval where reported, and the timescale over which degradation was observed.

Then, for each, answer one question explicitly: was the measurement of individuals aggregated, or of a collective capability such as a team's error-catching rate, a department's ability to handle a novel case, or an organization's performance during a system outage? Flag every study in the second category. That distinction is the one I care most about, and I expect the second category to be small.
Follow up in the same thread
  • Ask specifically for studies using outages, downtime or system failures as a natural experiment.
  • Ask what happened in organizations that reversed an automation decision.
P07SUPPORTING

Automation bias at the crew and team level

Named in the ledger as still to acquire. Gives the historical section teeth it currently lacks.

Paste this
Give me the classic and current literature on automation bias, automation complacency and appropriate reliance, focused on findings about teams, crews and organizational decision processes rather than single operators.

Start from Parasuraman, R. and Riley, V. (1997), "Humans and Automation: Use, Misuse, Disuse, Abuse", Human Factors, and Skitka, L., Mosier, K. and Burdick, M. (1999) on automation bias in cockpit tasks. Then bring the line forward to 2020 through 2026, including AI-specific work and the appropriate-reliance research programme (Schemmer and colleagues, and related work at CHI and CSCW).

For each: citation, DOI, design, sample, the finding, and effect size where reported.

Then answer one question directly, citing the evidence: does the literature show that crew or team structures amplify or dampen automation bias compared with individuals working alone? If the evidence is mixed or absent, say that, and describe what study would settle it.
Follow up in the same thread
  • Ask for the aviation and maritime incident-investigation literature specifically, not the lab studies.
  • Ask what training interventions have been tested against automation bias and whether any worked.
P08CORE

Redistribution of expertise after adoption

Dimension 2 of the instrument. Q3 asks whether people consult a colleague before, after, or instead of the tool. Find out whether anyone has measured that.

Paste this
Find empirical research on how adopting AI or algorithmic tools changes who is consulted, who is treated as an expert, and how status and authority are distributed inside an organization.

Include field studies, workplace ethnographies, quasi-experiments and large-scale observational studies of real workplaces.

Angles I want covered:
- Differential effects on junior versus senior workers
- What happens to the informal "go-to person" in a team
- Changes in referral, escalation and consultation patterns
- Deprofessionalization and the reallocation of discretion
- Any study that measured consultation sequence, meaning whether people ask a colleague before or after asking the tool

For each: author, year, venue, DOI, method, sample, setting, and the specific finding about redistribution.

Separate the results into two lists: studies that measured behavior, and studies that measured perception or self-report. Tell me how large each list is.
Follow up in the same thread
  • Ask for the Stack Overflow and internal-help-channel volume literature. Consultation traces exist there.
  • Ask what the medical literature says about changes in specialist referral after decision-support adoption.
P09CORE

Standards of evidence drifting without a decision

Dimension 1 of the instrument. The mechanism claim is that standards move through use rather than through policy.

Paste this
Find research documenting how introducing a tool changed what an organization accepts as sufficient evidence, sufficient documentation or sufficient justification for a decision, in the absence of any explicit policy change.

Likely locations for this literature: electronic health records and documentation quality, including note bloat and copy-forward; algorithmic risk assessment in criminal justice and child welfare; credit and underwriting automation; audit and assurance practice; academic peer review; and compliance functions.

The mechanism I am after is drift through use rather than change through decision. Prioritize accordingly.

For each source: author, year, venue, DOI, method, and the specific documented drift with whatever quantification exists.

Then identify separately any study that measured the same organization before and after adoption against the same criterion, since those are the only ones that can establish drift rather than difference. If there are none, say so.
Follow up in the same thread
  • Ask for the copy-forward and note-bloat quantitative literature by itself. It is unusually well measured.
  • Ask whether any regulator has documented standards drift in its own supervised institutions.
P10SUPPORTING

Who still measures unaided performance

Dimension 3 of the instrument, and the Budzyn problem. Nobody noticed the decline because in routine practice nobody works unaided.

Paste this
I want to know which professional fields retain any measurement of practitioner capability without tool assistance, and how they obtain it.

Find literature on:
- Aviation manual flying proficiency requirements and the evidence base behind them
- Anaesthesia, surgery and simulation-based assessment of unassisted performance
- Radiology reader studies conducted without computer-aided detection
- Skill degradation during periods of non-practice, in any field
- Any published protocol for periodically measuring a professional's unassisted performance

For each: citation, DOI, the protocol used, how frequently it is run, who mandates it, and what evidence exists that it actually detects degradation.

Finish with an assessment: which of these protocols could plausibly transfer to knowledge work, where there is no simulator, no licensing body and no natural task boundary? Be specific about what would break in the transfer.
Follow up in the same thread
  • Ask about the evidence behind mandated manual flying hours specifically. Is it empirical or regulatory folklore?
  • Ask whether any software organization runs deliberate no-assistance exercises and whether anything was published.
3
Track 3

Can it be measured?

The instrument is the whole bet. A framework paper without it collapses back to a framework note. The central methodological problem is that self-report is close to the wrong instrument for a change people by hypothesis cannot perceive.

P11CORE

Validated scales to anchor against

Point 3 of the instrument's validation plan. Face validity alone will draw a methods objection.

Paste this
I am building a short organizational survey instrument, roughly ten items across four subscales, administered at team level. I need existing validated scales to anchor construct validity against, meaning scales my instrument should correlate with meaningfully without being identical to.

Find validated instruments in organizational research measuring:
- Team psychological safety (Edmondson 1999 and later short forms)
- Organizational learning capability
- Team reflexivity
- Absorptive capacity
- Technology acceptance and post-adoption system use
- Decision quality and decision process quality
- Any existing scale measuring AI reliance, AI trust or human-AI team performance at group rather than individual level

For each: originating citation with DOI, number of items, reported reliability (Cronbach's alpha or omega), the population it was validated on, whether the items are freely reproducible, and a link to the item list if one is public.

Then recommend the two best candidates for a discriminant validity argument, meaning demonstrating that a new instrument is related to but distinct from an existing measure. Explain the reasoning for each choice, including what correlation range would support the argument and what range would sink it.
Follow up in the same thread
  • Ask what correlation with psychological safety would indicate the new scale is a psychological safety scale in disguise.
  • Ask for short forms only, four to six items. Response rate is the binding constraint.
P12CORE

Measuring a change people cannot perceive

The central design problem in the instrument draft. Asking people whether they over-rely measures self-image.

Paste this
My core methodological problem: I want to measure a change inside an organization that its members are, by hypothesis, unable to perceive. This makes conventional self-report close to the wrong instrument, because asking people whether they rely on a tool too much measures self-image rather than practice.

Find methods literature addressing this problem. Cover:
- Response shift bias and the retrospective pretest-posttest, also called the then-test
- Recall bias for one's own former attitudes, standards and competence, and its documented direction
- Document-anchored or artifact-based elicitation, where respondents react to a real past work product rather than to a memory
- Critical incident technique
- Cognitive interviewing for survey item development
- Behavioral trace and process-mining approaches as an alternative to self-report

For each: citation with DOI, what the method does, its documented failure modes, the sample sizes it typically requires, and its suitability for organizational settings as opposed to clinical or educational ones.

Finish with a recommended combination of two or three methods for a small-sample organizational study, roughly five to fifteen teams, and give the reasoning. Note explicitly where the recommended combination remains vulnerable.
Follow up in the same thread
  • Ask for evidence on whether the then-test actually corrects response shift or introduces its own bias.
  • Ask what an artifact-anchored protocol looks like operationally, step by step.
P13SUPPORTING

Behavioral traces as a reliance measure

Point 4 of the validation plan. One observable indicator beats ten self-reports, if it can be instrumented cheaply.

Paste this
Find research that measures reliance on AI or algorithmic systems using observable behavioral traces rather than self-report.

The kinds of measure I mean: override and deviation rates against a recommendation; time elapsed from recommendation to acceptance; edit distance between generated output and the output actually shipped; frequency of verification against an external source; rate of consultation of a human expert before or after the tool; and rate of documented disagreement.

For each study: author, year, venue, DOI, the exact measure used, how it was instrumented technically, the setting, the sample, and any validation of the measure against an independent outcome.

Two priorities. First, real workplaces over laboratory tasks; label which each study is. Second, instrumentation light enough that a small consulting engagement could reproduce it without engineering support. Rate each measure on that second criterion and say what it would take to collect.
Follow up in the same thread
  • Ask which of these measures have known gaming or Hawthorne problems once people know they are collected.
  • Ask for the clinical decision support override-rate literature specifically. It is the most mature version of this.
4
Track 4

How tools shape culture

The mechanism question, asked directly. This is where a position paper either has a theory of change or has a mood. Two prompts: the theoretical toolkit, and the documented cases that show the toolkit doing real work.

P14CORE

The theoretical toolkit for tools changing norms

The mechanism claim needs a named theory behind it. Ferdman's capacity-hostile environments already does the key move of relocating the unit of analysis from user to environment.

Paste this
Give me the strongest theoretical accounts of how a tool or technology changes the norms and practices of the group that adopts it, without anyone having decided to change them.

Cover this canon, and tell me explicitly what I am missing from it:
- Affordance theory, from Gibson through Norman to Hutchby's sociological version
- Script theory and the delegation of morality to artifacts (Akrich; Latour, "Where Are the Missing Masses?")
- Winner, "Do Artifacts Have Politics?" and the published rebuttals to it
- Infrastructure studies and invisible work (Star and Ruhleder; Bowker and Star)
- Domestication theory (Silverstone and Haddon)
- Sociomateriality and entanglement (Orlikowski; Barad) and the critiques of it as unfalsifiable
- Imbrication (Leonardi)
- Technological momentum (Hughes)
- Capacity-hostile environments (Ferdman, AI and Society, 2026)

For each: canonical citation with DOI, the specific causal mechanism the theory proposes in one sentence, what it explains well, and its main published critique.

Finish by naming the two or three that give the most usable mechanism for an empirical paper about organizational change following AI adoption, and say why. Usable means it generates a testable prediction, not just a vocabulary.
Follow up in the same thread
  • Ask which of these theories has ever been operationalized in a survey instrument.
  • Ask for the strongest published argument that all of this is unfalsifiable description.
P15CORE

Documented cases of a tool rewriting institutional norms

Precedent. If information systems have demonstrably rewritten institutional norms before, the AI case is a continuation with better instrumentation rather than a novel claim.

Paste this
Find well-documented empirical cases where introducing a specific tool or information system measurably changed an organization's norms, standards or internal distribution of authority. Before-and-after evidence strongly preferred.

Candidates to check and then go beyond:
- Electronic health records and the reshaping of clinical work, note-writing and the clinical gaze
- CompStat and policing practice
- Standardized testing and teaching practice
- Algorithmic scheduling in retail and hospitality
- Risk assessment instruments in criminal justice and child welfare
- Enterprise resource planning implementations
- Citation metrics and impact factors reshaping academic judgment

For each case: the single best scholarly source with DOI, the method, exactly what changed, over what timescale, and whether the change was intended by anyone at the time.

Sort the results into two groups and label them: studies with longitudinal or comparative designs, and single-site interpretive accounts. Tell me which of the longitudinal studies has the cleanest identification of the tool as the cause.
Follow up in the same thread
  • Ask for the citation-metrics case in depth. It is the one every academic reviewer has lived through.
  • Ask for cases where the norm change was later reversed, and what reversal required.
5
Track 5

The hostile reviewer

Run P16 against your own thesis monthly while drafting. Run P17 on every figure before it reaches a draft. These two prompts are the difference between a paper that survives review and one that gets a reviewer who has read the corrections you missed.

P16CORE

Steelman the opposite position

The decision brief already caught two overstatements this way. Vasconcelos turned the reflexive hook into a special case, and Vaccaro's creation-task gain crosses zero.

Paste this
Build the strongest evidence-based case AGAINST the following claim. Do not soften it, do not balance it, and do not add a conclusion restoring the original position. I want the opposition brief.

CLAIM: "Adopting AI tools degrades an organization's capacity for judgment in ways its members cannot perceive."

Include:
- Studies showing AI adoption improved organizational decision quality, learning or capability, with numbers
- Evidence that human-AI teams outperform either alone, and the specific conditions under which that holds
- Critiques of the deskilling literature as technologically determinist, as recycled moral panic, or as methodologically weak
- Historical cases where a predicted skill loss did not materialize, or was fully compensated for by a new capability
- Any failed replication, published comment or formal methodological critique of the studies most often cited for deskilling, specifically Bastani et al. on math tutoring, Kosmyna et al. "Your Brain on ChatGPT", and Budzyn et al. on endoscopy

For each: citation, DOI, and the specific claim it supports.

End with the three strongest sentences a hostile reviewer could write against the original claim. Write them as a reviewer would write them, not as a summary.
Follow up in the same thread
  • Run the same prompt against the specific construct: 'institutional interpretability is a distinction without a difference.'
  • Ask which of the counter-arguments has no good answer, and say so plainly.
P17EVERY TIME

Verify before it enters the ledger

The gate between a search result and the Evidence Ledger. Three corrections were already caught this way on August 18.

Paste this
Verify the following claim against its primary source and report every discrepancy precisely.

CLAIM AS I CURRENTLY HAVE IT: [PASTE THE CLAIM AND THE CITATION EXACTLY AS WRITTEN IN YOUR DRAFT]

Check and report on each of these separately:
1. The exact figures, including percentages, effect sizes and confidence intervals
2. Sample size and composition
3. Study design, and whether it supports a causal reading
4. Statistical significance, and whether any reported confidence interval crosses zero
5. Peer-reviewed or preprint, and the venue
6. Whether the paper has been retracted, corrected, or formally critiqued in a published comment
7. Whether the claim as I have written it is what the authors actually concluded

Quote the relevant sentence from the source verbatim, with a page or section reference.

If my claim overstates the source in any respect, say so directly and then give me the strongest version of the claim that the source does support.
Follow up in the same thread
  • For any preprint: ask whether a peer-reviewed version now exists and whether the numbers changed.
  • For any contested study: ask for the published comment and the authors' reply, both.
6
Track 6

Where it lands

Positioning, and the Pittsfield event. P20 matters early rather than late: whether the event doubles as the instrument's first deployment changes the event's design, and that decision has to be made before it is scheduled.

P18SUPPORTING

Venues and who is already in the room

Different venues reward different shapes. A taxonomy without data gets desk-rejected at CHI and is publishable at Big Data and Society.

Paste this
I am writing a position paper arguing that the consequential interpretability problem is institutional rather than mechanistic: whether an organization can perceive how a tool changed its norms, its standards of evidence and its distribution of expertise. The paper comes with a short original survey instrument.

Give me three things.

1. Venues. Assess fit for: FAccT, CSCW, CHI, Big Data and Society, Science Technology and Human Values, AI and Society, Organization Science, Journal of Management Studies, Learning Media and Technology, and any I have not named that fit better. For each: typical length, review cycle and acceptance rate if published, one sentence on what that venue rewards, and one sentence on what reliably gets rejected there.

2. People. The five to ten researchers publishing closest to this argument right now, each with their single most relevant recent paper and DOI. Say for each whether they are a likely reviewer, a likely citation, or a likely competitor.

3. Calls. Any active special issue or call for papers on AI and organizational change, AI and expertise, or AI governance, closing within the next twelve months, with deadline and link.

Be concrete about deadlines and do not invent them. If you cannot confirm a call is still open, label it unconfirmed.
Follow up in the same thread
  • Ask what a FAccT reviewer does to a paper whose empirical contribution is a ten-item survey.
  • Ask which of these venues has published anything with 'legibility' in the title.
P19SUPPORTING

What the standards bodies already call this

Determines whether the construct fills a gap in AI governance or duplicates a control that already exists on paper.

Paste this
Map the practitioner and governance landscape for organizational AI oversight, so I can see what already exists before proposing anything new.

Cover:
- The NIST AI Risk Management Framework, particularly the Govern function and its subcategories
- ISO/IEC 42001 and what an AI management system actually requires
- The EU AI Act's organizational obligations, including AI literacy requirements under Article 4
- Internal audit and assurance frameworks for AI (IIA, ISACA and comparable)
- AI maturity and readiness models published by major consultancies and standards bodies

For each: what it actually requires or measures at the organizational level, in specific terms rather than in its own marketing language.

Then answer one question directly for each: does it assess an organization's awareness of how its own practices, standards or expertise distribution have changed since adoption, or does it only assess process compliance and documented control?

Label each item as an enforceable legal requirement, a voluntary certifiable standard, a professional framework, or marketing. Be blunt about which is which. End by naming the single clearest gap.
Follow up in the same thread
  • Ask whether any framework requires measuring staff capability without the system.
  • Ask what an ISO 42001 auditor actually asks for as evidence, at the level of documents requested.
P20CORE

The public event as legitimate research

Open question 3 in the instrument draft. Arielle's IRB certification makes this feasible, and Akgun 2026 is the citable precedent.

Paste this
Find precedents and methods for running a public community event about AI that also functions as legitimate research rather than as outreach.

Cover:
- Participatory and community-based design research methods, with their methodological credentials
- Deliberative formats used for technology policy: citizens' juries, consensus conferences, deliberative polling, citizen panels
- The public engagement with science and technology literature on evaluating such events
- Practical research ethics: informed consent in a public setting, IRB or ethics-board treatment of public events, and data handling for recorded public conversation

Start from and go beyond: Akgun, S., Choi, K. and Lee, H. R. (2026), "Designing a critical AI literacy program for K-8 STEM education: adopting a community-centered approach", Disciplinary and Interdisciplinary Science Education Research, DOI 10.1186/s43031-026-00155-1.

For each source: citation with DOI, the format used, the size and setting of the event, what data was collected, how consent was handled, and what was published from it.

Finish by recommending the two or three formats most workable for a single evening event in a small city, roughly thirty to eighty attendees, one facilitator and one note-taker, where the goal is both a genuine public conversation and usable research data. State what each format demands that a standard community event does not.
Follow up in the same thread
  • Ask what consent language is standard for recording a public deliberative event and publishing from it.
  • Ask what these events produce that a survey cannot, and be specific.

What to do with the results

01

Run Track 0 first, in one sitting.

Three prompts. Any strong hit changes what the article is. If P01 and P02 both come back genuinely empty, the territory is open and the rest of the pack is worth the time.

02

Everything lands in the ledger, or it does not exist.

The Evidence Ledger at 01_Papers/AI Interpretability Paper/02_Evidence/Evidence_Ledger.md is the single record. Search results that never make it there will be re-found and re-checked in three weeks.

03

Verify with P17 before writing anything into a draft.

Every figure, every sample size, every confidence interval. The three corrections already caught on August 18 were all of this kind.

04

Sort into the four dimensions as you go.

Moved standard of evidence, redistributed expertise, lost unaided baseline, explanation as terminal. Sources that fit no dimension are telling you either that the dimension is missing or that the source is off-topic. Both are worth noticing.

05

Track what came back empty.

Keep a short list of the searches that found nothing. In a position paper the documented absence of literature is a claim you can make, and it is only credible if you can say what you searched for.