Institutional Legibility Instrument, v0
1 of 6. The constructthe degree to which a group can perceive how a tool has changed its own norms, standards of evidence, and distribution of expertise.
Institutional legibility: the degree to which a group can perceive how a tool has changed its own norms, standards of evidence, and distribution of expertise.
Low institutional legibility does not mean a team uses AI badly. It means the team cannot see what changed. A team could have adapted well and still score low, and that is the point. The construct is about perception of change, not quality of outcome. Keep these separate in every item or the instrument measures satisfaction instead.
2 of 6. The core design problemSelf-report is close to the wrong instrument for an unperceived change.
Self-report is close to the wrong instrument for an unperceived change. The entire claim is that people cannot see this happening to them, so asking "do you rely on AI too much?" measures self-image.
The rule for every item: ask about observable practice, never about perception.
| Do not ask | Ask instead |
|---|---|
| Do you rely on AI too much? | When did this team last review a decision where the tool's recommendation was accepted without discussion? |
| Has AI changed your standards? | Describe a piece of work this team would have sent back a year ago and would accept today. |
| Do you trust the AI? | Who gets consulted before the tool, and who gets consulted after it? |
| Are you still good at this? | When did anyone here last do this task unaided, and how did it go? |
The fourth row is Budzyń. Nineteen expert endoscopists lost 6 percentage points of unaided detection and nobody noticed, because in routine practice nobody works unaided any more. Every field has an equivalent, and the item that finds it is the one that matters.
3 of 6. Four dimensions, with candidate itemsTen items total is the target.
Ten items total is the target. Ten items get completed. Thirty items get abandoned at question eleven and produce a biased sample of the conscientious.
1. Moved standard of evidence
What the team now accepts as sufficient, which it would previously have questioned, without anyone having decided to change.
- Q1. Think of the last piece of work this team approved that was substantially AI-assisted. What was checked before approval, and by whom? (Open response, coded.)
- Q2. In the past three months, how many times has a generated rationale or summary ended a discussion rather than started one? (Never / Once or twice / Monthly / Weekly or more / I would not be able to tell.)
The final option on Q2 is the most informative response available and it should appear on several items. Someone who cannot tell is reporting low institutional legibility directly.
2. Redistributed expertise
Who gets asked, and where in the sequence.
- Q3. When someone here has a question in your area of expertise, do they typically come to you before consulting an AI tool, after, or instead? (Before / After / Instead / Varies / Do not know.)
- Q4. Name a person on this team whose judgment was routinely sought eighteen months ago and is sought less now. (Open, optional, anonymized in analysis.)
Q4 is the sharpest item in the draft and the most likely to be refused. Test whether an anonymous format rescues it. If it does not, replace it with a role-level version rather than dropping the dimension.
3. Lost unaided baseline
Whether the group retains any measurement of its own capability without the tool.
- Q5. When did anyone on this team last complete a core task in your work without AI assistance? (Within the past week / Month / Quarter / Year / Longer / Do not know.)
- Q6. If the tool were unavailable for two weeks, which specific tasks would take materially longer, and which would come out materially worse? (Open, two-part.)
Q6 separates speed loss from quality loss. Teams reliably predict the first and reliably fail to predict the second, and the gap between those two answers is itself a measure.
4. Explanation as terminal
Whether interpretability output opens inquiry or closes it. This is where Jenkins, Bansal and Vasconcelos all land.
- Q7. When a tool provides a rationale for its output, what typically happens next? (Someone verifies against a source / Someone asks a follow-up of the tool / It is accepted and work continues / It is ignored entirely / Varies.)
- Q8. Has anyone here overruled an AI output in the past month? What happened to that decision afterwards? (Open.)
Cross-cutting
- Q9. How did this team decide to adopt the tools it uses? (Deliberate evaluation / A pilot that became permanent / Individuals adopted separately and it spread / Mandated from above / Do not know.)
- Q10. What is one thing this team does differently now than eighteen months ago that nobody explicitly decided to change? (Open.)
Q10 is the whole instrument in one question. If the ten-item version has to shrink, this is the item that survives.
4 of 6. Scoring approachDo not build a single composite score.
Do not build a single composite score. A composite invites reading it as a grade and the construct is not evaluative.
Report four subscores, one per dimension, plus a separate opacity count: the number of items answered "do not know" or "would not be able to tell." The opacity count may end up being the most predictive quantity in the instrument, and it is the one that most directly operationalizes the construct.
5 of 6. Validation plan1.
- Face validity. Six to ten cognitive interviews. Read each item aloud, ask what the respondent thinks it is asking. Rewrite anything that gets a different answer twice.
- Discrimination. The instrument has to separate teams. If every team scores in the same band, the items are measuring the era rather than the team.
- Construct anchoring. Correlate against an existing validated measure. Team psychological safety (Edmondson) is the obvious candidate: it is well validated, it is plausibly related, and it should be related without being identical. If the correlation is near 1.0 the instrument is a psychological safety scale wearing a hat.
- Behavioral proxy, if reachable. One observable indicator beats ten self-reports. Frequency of documented overrides, or time from AI output to approval, if any client instruments their workflow.
6 of 6. Open questions for the working session1.
- Unit of response. One respondent per team, or every member? Every member is better and costs the response rate. A hybrid, with one long form for a lead plus a three-item version for everyone else, is worth pricing out.
- Retrospective anchoring is fragile. Several items ask about eighteen months ago. Memory for one's own former standards is poor and systematically biased toward continuity, which biases the instrument toward finding no change. Is there a version that anchors on a documented artifact instead, such as a work product from before adoption that the team re-reviews?
- Does the community event supply the first deployment? Arielle's IRB certification makes this feasible. Akgün et al. (2026) is the citable precedent for treating a public participatory event as design research rather than outreach. Deciding this changes the event's design, so it should be decided before the event is scheduled rather than after.
- Does it double as a billable FWC diagnostic? If yes, there are two versions: the research instrument and the client-facing report. The client version needs a constructive frame, and the research version has to stay free of anything that would bias a respondent toward the flattering answer. Building both from the start is easier than retrofitting either.
- Who owns it? Whoever owns the instrument owns the paper's empirical claim, carries the methods burden, and has a strong argument for first authorship. Settle this before item drafting rather than after.
Sections in this document
- 1. The construct
- 2. The core design problem
- 3. Four dimensions, with candidate items
- 4. Scoring approach
- 5. Validation plan
- 6. Open questions for the working session