Three legibilities: model, interaction, institutionsource .md1 sections

Three legibilities: model, interaction, institution

Palmer Foote, August 2026. The memo that separates the three jobs the word "interpretability" is doing. The app is an instrument for the second and third. Source text reproduced as written.

1 of 1. The memoGenuinely interesting question, because the word "interpretability" is doing two different jobs and the gap between them is where your opening is.

Genuinely interesting question, because the word "interpretability" is doing two different jobs and the gap between them is where your opening is.

The field means something narrower than Jenkins does. Mechanistic interpretability studies the artifact — features, circuits, sparse autoencoders, attribution graphs, and increasingly chain-of-thought monitoring. Its object is the model's internals, its audience is researchers, and its motivating question is safety. MIT Tech Review named it a 2026 breakthrough technology, and the work is concentrated in a handful of labs (Anthropic, OpenAI, DeepMind, with Neuronpedia as the main open-access effort). Jenkins's transparency problem studies the relationship — whether a person can articulate how a system is shaping their perception. Different object, different audience. Borrowing one word for both is a metaphor, and if IM's product story ever lets "the user can see how the layer shapes them" blur into "we can trace the circuit," that's an integrity problem waiting to happen.

But four real intersections:

Interpretability outputs are media. The moment a feature dashboard or an attribution graph is shown to someone who isn't a researcher, it becomes a representation — with genre conventions, aesthetic choices, and things it leaves out. Jenkins's sharpest empirical finding is that students playing a history game absorbed the game's account of history as authentic; they lacked "a vocabulary to critique how the game itself constructed history." That is precisely the failure mode for interpretability UIs. An explanation that looks legible invites people to mistake it for the model. XAI research has its own name for this — explanation-induced overreliance — and it arrives at Jenkins's conclusion from the other direction.

Chain-of-thought is the acute case. It's the most human-readable interpretability artifact and the one with the liveliest open question about whether stated reasoning is faithful to actual computation. Most readable, most at risk of being read as authentic. Jenkins's exact structure.

Interpretability is the missing chapter of connectivism. Siemens said learning resides in non-human appliances and that you should connect to the node and let the flow work. He has no account whatsoever of auditing the node. Mech interp is the first serious attempt to read what the appliance actually knows — which means the 2005 paper's central claim now has an empirical research program attached to it that Siemens never anticipated. Worth saying out loud today; it's the cleanest bridge between the two documents and the technical field.

The participation gap reproduces itself one level up. Very few people can see inside these models. Everyone else receives whatever explanation is passed down. Jenkins's core move — the divide isn't hardware access, it's who finds the environment rich versus thin — transfers without modification.

The frame:

  1. Model interpretability — what's happening in the weights. Not your fight, and you don't need it to be.

  2. Interaction interpretability — can this user tell how the system shaped this output. Contested, over-trust risk, active HCI field.

  3. Institutional interpretability — can a team see how the tool has changed its norms, its judgment, its distribution of expertise. Nearly vacant. That's Jenkins and Siemens territory, and it's FWC's turf.

Intelligence Meridian's "editable cognitive layer" sits at 2 and 3, downstream of 1 rather than competing with it. That's a defensible position, and it's honest about what you can and can't claim.

One caution for credibility: even the technical field's own artifacts are lossy — recent SAE work notes plainly that these methods "do not fully solve polysemanticity," meaning features remain partly entangled. Don't let anyone in the room, including yourself, treat interpretability as solved at any level.

Sections in this document