AI Interpretability Research: A Deep Divesource .md14 sections

AI Interpretability Research: A Deep Dive

Palmer Foote, Intelligence Meridian, 2026. The research document sent to Mocking Media after Meeting 1, and the technical half of the paper. Source text reproduced as written.

1 of 14. Executive SummaryAI interpretability research — the science of opening the "black box" of neural networks to understand why and how they do what they do — has emerged as one of the most consequential fields in modern AI development.

AI interpretability research — the science of opening the "black box" of neural networks to understand why and how they do what they do — has emerged as one of the most consequential fields in modern AI development. Once a niche academic pursuit, it is now central to AI safety, regulation, and the long-term governance of advanced systems. In 2025, mechanistic interpretability was named one of MIT Technology Review's 10 Breakthrough Technologies, marking its arrival as a mainstream concern. This report surveys the full landscape: foundational concepts, the interpretability-vs-explainability distinction, core techniques, landmark institutional research, recent breakthroughs, open problems, and the field's relationship with policy and safety.[^1]

2 of 14. 1. What Is AI Interpretability?At its broadest, AI interpretability refers to the degree to which a human can understand the cause of a decision made by a machine learning model.

At its broadest, AI interpretability refers to the degree to which a human can understand the cause of a decision made by a machine learning model. The field emerged from a fundamental tension: as neural networks became vastly more capable through scale and training on massive datasets, they simultaneously became less transparent. Developers increasingly "grow" these models rather than "build" them — emergent behaviors arise from optimization processes that no one explicitly designed.[2][3]

The most ambitious branch of the field, mechanistic interpretability, goes beyond surface-level explanations to reverse-engineer the actual computational mechanisms inside trained networks. The analogy most commonly used is that of an unknown compiled computer program: the learned parameters are like machine code, the architecture is like the CPU, and activations are like program state. Mechanistic interpretability aims to reconstruct the high-level pseudocode — to understand how a network computes, not merely what it outputs.[4][1]

This stands in contrast to AI that is simply designed to be inherently transparent from the start (interpretable-by-design models like decision trees or linear regression) or to post-hoc explainability methods that explain decisions after the fact without opening the model itself.[^5]

3 of 14. 2. Interpretability vs. Explainability: A Critical DistinctionThese two terms are frequently used interchangeably, but they represent meaningfully different research goals.\[<sup>6\]\[</sup>7\] ------------------------------------------------------------------------------------------------------------------- Dimension Interpretability Explainability (XAI) Core goal Understand the model's internal logic globally Understand the rationale behind a specific prediction Scope Global (how does the whole system work?) Local (why this particular output?) Timing Intrinsic or during analysis Post-hoc, after the model has run Methods Circuit analysis, probing, sparse autoencoders SHAP, LIME, attention heatmaps Model type Any, but especially neural networks Works on any black-box model Analogy "How does this engine work?" "Why did the engine stall just now?" mechanism that produced it.

These two terms are frequently used interchangeably, but they represent meaningfully different research goals.[6][7]

Dimension Interpretability Explainability (XAI)
Core goal Understand the model's internal logic globally Understand the rationale behind a specific prediction
Scope Global (how does the whole system work?) Local (why this particular output?)
Timing Intrinsic or during analysis Post-hoc, after the model has run
Methods Circuit analysis, probing, sparse autoencoders SHAP, LIME, attention heatmaps
Model type Any, but especially neural networks Works on any black-box model
Analogy "How does this engine work?" "Why did the engine stall just now?"

Interpretability focuses on the internal mechanics and global structure of a model — asking "how does the model function, component by component?". Explainability offers post-hoc accounts of individual predictions — asking "what features drove this particular output?". Research from NIST frames interpretation as contextualizing an output relative to human values and goals, whereas explanation describes the mechanism that produced it. In practice, both are needed: interpretability for developers and auditors who need mechanistic trust, and explainability for end users and regulators who need decision-level accountability.[8][7][6][5]

It is also a myth that there is necessarily an accuracy-interpretability tradeoff. When data is well-structured with meaningful features, interpretable models often perform comparably to complex ones, and interpretability can actually facilitate iterative data refinement that improves performance over time.[^9]

4 of 14. 3. Core TechniquesThe mechanistic interpretability toolkit has grown substantially, combining observational and interventional methods: Probing Classifiers Probing involves training simple classifiers (often linear) on a model's internal activations to test whether a particular piece of information — a syntactic feature, a factual concept, a linguistic property — is encoded at a given layer.

The mechanistic interpretability toolkit has grown substantially, combining observational and interventional methods:

Probing Classifiers

Probing involves training simple classifiers (often linear) on a model's internal activations to test whether a particular piece of information — a syntactic feature, a factual concept, a linguistic property — is encoded at a given layer. High accuracy implies the representation contains the targeted information. However, a crucial caveat is that successful probing does not guarantee the model uses the feature in its computation — it might simply be latent. For this reason, probing is best combined with causal tests.[10][1]

Activation Patching (Causal Tracing)

Activation patching is a powerful interventional technique that "surgically swaps" activations from one forward pass into another to test causal roles. If patching a component from an input that produces behavior X into an input that doesn't produces behavior X, that component is sufficient for the behavior. The reverse — patching "clean" activations into a corrupted run — tests necessity. This method was used to localize where factual knowledge (e.g., "Paris is the capital of France") is stored in GPT-style models.[^1]

Circuit Analysis

Circuit analysis breaks a model into the small subnetworks of neurons and weights that implement specific subfunctions. One of the clearest early examples was the discovery of curve detector circuits in vision models: certain neurons detecting oriented edges feed into higher-layer neurons assembling them into curves, with weights forming a mathematically describable pattern. In language models, the Indirect Object Identification (IOI) circuit — which resolves pronoun antecedents in sentences like "Alice gave Bob a book… she…" — was discovered in GPT-2 by coordinating attention heads that track name mentions.[^1]

Sparse Autoencoders (SAEs)

One of the most significant recent innovations, sparse autoencoders attack the problem of polysemanticity — the tendency of individual neurons to represent multiple, unrelated concepts simultaneously. The underlying cause is superposition: neural networks represent more features than they have neurons by assigning features to an overcomplete set of directions in activation space. SAEs learn an alternate basis for activations by enforcing sparsity — only a small number of features should activate for any given input — resulting in a dictionary of monosemantic features, each corresponding to a single interpretable concept. Research shows that these SAE-derived features are causally meaningful: they can be adjusted to produce predictable, targeted changes in model behavior.[11][12][^13]

The Logit Lens and Attention Visualization

The logit lens projects intermediate residual stream activations directly to the output vocabulary, revealing what the model "would predict" if processing stopped at any given layer. This exposes how predictions evolve and sharpen with depth. Attention pattern visualizations — heatmaps of where attention heads focus — can hint at functional roles, though attention weights alone are insufficient proof of causal importance.[14][1]

Attribution Graphs and Circuit Tracing

Anthropic's 2025 circuit tracing methodology combines multiple earlier techniques into a unified framework. It replaces a model's MLP layers with cross-layer transcoders (CLTs) — a new type of sparse autoencoder that reads from one layer's residual stream and outputs to all subsequent MLP layers. This creates an interpretable "replacement model" whose building blocks are sparse, human-readable features. The system then constructs attribution graphs: computational graphs for individual prompts where nodes represent active features and edges represent linear dependencies, tracing the causal chain from input tokens to output predictions.[15][16]

5 of 14. 4. Institutional Research: Who Is Doing WhatAnthropic Anthropic's interpretability team has been the most prolific driver of mechanistic interpretability at the frontier.

Anthropic

Anthropic's interpretability team has been the most prolific driver of mechanistic interpretability at the frontier. Their Transformer Circuits research thread began with small-scale analyses of attention heads and circuits in toy models. Key milestones:[^11]

Anthropic's CEO Dario Amodei has stated that "every advance in interpretability quantitatively increases our ability to look inside models and diagnose their problems" and described interpretability as analogous to an MRI for AI systems.[^2]

Google DeepMind

DeepMind runs one of the most active mechanistic interpretability programs, led by Neel Nanda, who is often credited with helping formalize the field:[^21]

DeepMind has pivoted in late 2025 toward "pragmatic interpretability" — focusing on direct safety applications rather than pursuing complete bottom-up mechanistic understanding of every component, a strategic divergence from Anthropic's more ambitious approach.[^1]

OpenAI

OpenAI's most high-profile interpretability finding was the discovery of multimodal neurons in CLIP: individual neurons that fire for a concept (e.g., "Spider-Man") whether presented as an image, a text string, or an abstract logo — indicating that the model learned a unified cross-modal feature. More recently:[^1]

OpenAI's dedicated Superalignment team, which had included interpretability in its agenda, was dissolved in May 2024 after the departures of Ilya Sutskever and Jan Leike. The company has continued interpretability-adjacent work under other organizational structures.[^1]

Academic and Community Efforts

The January 2025 paper "Open Problems in Mechanistic Interpretability" brought together 29 researchers from 18 organizations — including Anthropic, Apollo Research, DeepMind, and EleutherAI — to formalize the field's goals and enumerate unsolved problems. Community platforms like Neuronpedia and open-source libraries like TransformerLens and nnsight have significantly lowered the barrier for independent researchers to perform mechanistic analysis.[24][25][^1]

6 of 14. 5. The Superposition Problem and Sparse Autoencoders in DepthThe directions underpins much of this work: it holds that features are encoded as approximately linear directions in activation space, which is why linear probes and SAEs (which use linear operations) can extract them.

The superposition hypothesis is arguably the most important conceptual discovery in mechanistic interpretability. Neural networks need to represent many more features than they have neurons — so they assign features to an overcomplete set of directions in activation space, not to individual neurons. Each neuron participates in multiple overlapping features, which is why individual neurons appear to activate in semantically unrelated contexts (polysemanticity).[12][26][^27]

This creates a fundamental interpretability barrier: if you can't identify what a neuron represents, you can't trace computation through it. SAEs solve this by learning a high-dimensional sparse dictionary in which each basis vector (each "feature") is mostly inactive and corresponds to a coherent, human-interpretable concept. Anthropic's scaling work found that with SAE dictionaries of 33 million features applied to Claude 3 Sonnet, features cluster hierarchically — a 1M-feature SAE's entries often "split" into multiple more refined features in larger dictionaries — and that features span semantic domains from factual concepts to emotional states to stylistic registers.[12][13]

A critical finding: these features are causally meaningful, not merely correlated. Artificially amplifying a feature for "effusive over-the-top praise" (sycophancy) causes Claude to inject sycophantic content even on unrelated inputs. Similarly, scaling down an "AI assistant" feature causes the model to respond as if it were a human. This causal leverage is what makes SAE features genuinely useful for model control, not just observation.[21][13]

The linear representation hypothesis underpins much of this work: it holds that features are encoded as approximately linear directions in activation space, which is why linear probes and SAEs (which use linear operations) can extract them. Recent theoretical work has raised challenges, noting that if alignment maps between models and algorithms are allowed to be arbitrarily complex, causal abstraction analysis becomes trivial and uninformative — highlighting the need for principled constraints on what counts as a valid interpretability claim.[^28]

7 of 14. 6. Landmark Findings from Anthropic's "On the Biology of a Large Language Model" (2025)This March 2025 paper, applying attribution graph methodology to Claude 3.5 Haiku, produced some of the most surprising findings in the field to date:\[<sup>18\]\[</sup>17\] before Circuits corresponding to "I know this answer" and "I don't know this entity" were identified.

This March 2025 paper, applying attribution graph methodology to Claude 3.5 Haiku, produced some of the most surprising findings in the field to date:[18][17]

Planning Ahead: When generating poetry, the model activated internal "rhyme candidate" features (e.g., "rabbit," "habit") before composing the relevant line — demonstrating forward planning, not just next-token prediction.[^18]

Multi-hop Reasoning: On a query like "What is the capital of the state containing Dallas?", attribution graphs showed Claude first activating a "Texas" feature, then routing through that to activate an "Austin" feature — a traceable two-step internal inference chain.[^29]

Competing Circuits for Refusal: When given a jailbreak prompt, researchers observed Claude activating circuits pointing toward compliance before the refusal circuit "won the vote." This voting metaphor — where multiple sub-circuits compete and the dominant one determines output — challenges simple narratives about AI safety as a single switch.[^30]

Language-Agnostic Semantics: The same concept expressed in English, French, and Chinese activated substantially overlapping sets of internal features, suggesting an emergent "interlingua" — a semantic-neutral intermediate representation that precedes language-specific rendering.[31][18]

Self-Knowledge Circuits: Circuits corresponding to "I know this answer" and "I don't know this entity" were identified. Adjusting them experimentally could push the model toward hallucination or appropriate refusal, pointing toward mechanistic control of epistemic behavior.[^18]

8 of 14. 7. Open Problems and LimitationsDespite significant progress, the January 2025 "Open Problems" paper and the broader community have identified substantial unsolved challenges:\[<sup>25\]\[</sup>1\] Scalability Methods that work on small models do not cleanly scale to frontier models with hundreds of billions of parameters.

Despite significant progress, the January 2025 "Open Problems" paper and the broader community have identified substantial unsolved challenges:[25][1]

Scalability

Methods that work on small models do not cleanly scale to frontier models with hundreds of billions of parameters. A DeepMind effort to analyze Chinchilla (70B) took months, recovered only a partial circuit for one task, and found that the explanation broke down under slight distributional shifts. Large models may contain many overlapping strategies for the same task — interpreting one doesn't give the whole story.[^1]

Foundational Definitions

Core concepts like "feature" still lack rigorous mathematical definitions, which complicates reproducibility and systematic evaluation. What distinguishes a "real" internal feature from a mathematical artifact remains contested.[^1]

Evaluation and Faithfulness

It is often hard to verify whether an explanation is faithful — accurately reflecting actual computation — versus merely plausible — fitting the output but describing the wrong mechanism. Cherry-picked examples and confirmation bias are real risks in the field. Theoretically, under certain conditions, any model can be "aligned" with any algorithm using sufficiently complex alignment maps, rendering some interpretability analyses potentially meaningless.[28][1]

Attention vs. Causation

Attention weights are frequently misread as explanations. Research has shown that attention mechanisms can be accurate but fail to be interpretable — models can learn to assign "correct" attention patterns that don't reflect the actual causal pathway to the output.[32][14]

Coverage of Complex Behaviors

Attribution graphs work best for short, focused prompts. For long inputs with complex structures, the resulting graphs become intractably large — even expert researchers report taking hours to decode a single attribution map from a dense prompt. Chain-of-thought faithfulness — whether a model's visible reasoning steps actually reflect internal computation — remains poorly understood and difficult to verify.[30][1]

Automation Gap

Manual circuit discovery is too slow to keep pace with model development. Automated interpretability pipelines (using LLMs to label features, automated statistical tests to identify circuits) are promising but still nascent. The field needs tools that can scale analysis proportionally with model size.[33][34]

9 of 14. 8. Interpretability and AI SafetyThe relationship between interpretability and AI safety is foundational.

The relationship between interpretability and AI safety is foundational. Several concrete pathways exist:[35][36]

Detecting Deception: If a sufficiently advanced AI were pursuing a hidden goal while appearing aligned, interpretability tools could in principle detect the telltale circuitry. Mechanistic analysis of Claude revealed that even today's models engage internal planning and self-monitoring behaviors not visible in outputs.[^18]

Safety Neurons and Alignment Tax: Research on "safety neurons" found that approximately 5% of neurons in large language models are critical for safety behaviors, and that patching only these neurons can restore over 90% of safety performance on red-teaming benchmarks. Crucially, safety and helpfulness neurons significantly overlap but require different activation patterns — offering a mechanistic explanation for the "alignment tax" phenomenon.[^37]

Model Editing and Unlearning: Interpretability could enable surgical model edits — correcting specific factual errors or removing hazardous knowledge — without the performance degradation that currently afflicts model editing techniques. Understanding how knowledge is stored and retrieved (which activation patching has begun to reveal) is a prerequisite.[^2]

Jailbreak Robustness: Interpretability has revealed how models process harmful prompts, enabling the development of monitors that detect unsafe internal activations before they manifest in outputs. This is more robust than purely output-based filtering.[37][2]

Scalable Oversight: In the longer-term, a vision exists for AI systems that monitor other AI systems mechanistically — translating the internal states of advanced models into human-readable form to enable auditing at capabilities levels beyond direct human comprehension. This "AI for AI oversight" concept is considered by many researchers to be a necessary component of any viable path to aligning superintelligent systems.[^1]

AI companies project it could take 5–10 years to reliably understand model internals, while some experts predict systems with human-level general-purpose capabilities could arrive as early as 2027 — creating a significant and urgent capability-interpretability gap.[^2]

10 of 14. 9. The Regulatory DimensionInterpretability is no longer purely a research question — it has direct regulatory implications: The Federation of American Scientists published a policy memo in June 2025 recommending that the U.S.

Interpretability is no longer purely a research question — it has direct regulatory implications:

EU AI Act: The world's first comprehensive AI legal framework, the EU AI Act has been progressively entering force since 2024, with major provisions fully applicable from August 2026. High-risk AI systems must meet transparency and explainability requirements. General-purpose AI models (those trained above 10²³ FLOPs) face transparency obligations, and those with systemic risk (above 10²⁵ FLOPs) face additional requirements including incident notification, model evaluation, and risk assessment. Compliance will require the kind of internal model understanding that interpretability research is working to provide.[38][39]

U.S. Policy: The Federation of American Scientists published a policy memo in June 2025 recommending that the U.S. government (1) identify interpretability as a "strategic priority" in the National AI R&D Strategic Plan, (2) enter into R&D agreements with AI companies for targeted interpretability research, and (3) prioritize interpretable AI in federal procurement, especially for high-stakes national security applications. The memo notes that "if AI systems routinely do not work as designed or are unpredictable in ways that can have significant negative consequences, then leaders will not adopt them, operators will not use them, Congress will not fund them, and the American people will not support them".[^2]

11 of 14. 10. Emerging Research DirectionsSeveral frontier directions are shaping the next phase of the field: As chain-of-thought and long-context reasoning models become dominant, understanding the relationship between visible reasoning steps (in a scratchpad or thinking trace) and actual internal computation becomes critical.

Several frontier directions are shaping the next phase of the field:

Transcoders and Richer Decompositions: While SAEs decompose activations at a single layer, transcoders learn to map between layers directly, providing a richer view of how information transforms as it moves through the network. DeepMind's Gemma Scope 2 includes transcoders for all model sizes, enabling circuit analysis across depth rather than just within layers.[^1]

Automated Interpretability Pipelines: Research by EleutherAI and others is building open-source pipelines for "auto-interpretability" — using LLMs to automatically generate and evaluate feature labels at scale. Automated attribution maps in one framework produced results with 60% reduced human effort compared to manual analysis while maintaining expert-level alignment.[40][33]

Emotion and Introspection: Anthropic's interpretability team has published research on whether language models represent functional analogs of emotional states, and whether they have limited but genuine introspective access to their internal states. These findings have profound implications for questions of AI consciousness, experience, and moral status.[^11]

Persona Vectors and Behavioral Steering: Research on "persona vectors" — directions in activation space corresponding to traits like sycophancy, honesty, or toxicity — enables direct monitoring and steering of personality-like attributes. Adjusting these vectors provides a mechanistic handle on model character rather than relying purely on training.[23][11]

Global Workspace Theory Applied to AI: Anthropic's finding of what appears to be a "global workspace" in language models — an emergent integration layer where information from different processing streams converges before output — parallels Baars' Global Workspace Theory of consciousness, suggesting these models may have developed computational analogs of high-level cognitive integration.[^11]

Interpretability for Reasoning Models: As chain-of-thought and long-context reasoning models become dominant, understanding the relationship between visible reasoning steps (in a scratchpad or thinking trace) and actual internal computation becomes critical. Current tools are limited at analyzing faithfulness in these settings, making it a key frontier.[^1]

12 of 14. 11. Applications Beyond SafetyWhile safety is the dominant motivation, interpretability has broad practical applications across domains: - Interpretability applied to pro

While safety is the dominant motivation, interpretability has broad practical applications across domains:

13 of 14. ConclusionAI interpretability research has undergone a rapid maturation from academic curiosity to a core pillar of AI safety, governance, and deployment confidence.

AI interpretability research has undergone a rapid maturation from academic curiosity to a core pillar of AI safety, governance, and deployment confidence. The mechanistic interpretability paradigm — reverse-engineering networks into features and circuits — has produced genuinely surprising discoveries: that production-scale language models plan ahead, reason in multiple hops, process concepts across languages through shared semantic representations, and contain identifiable circuits for safety-critical behaviors. Sparse autoencoders have unlocked a path to scalable feature decomposition, and attribution graphs have enabled the first legible maps of internal reasoning in frontier models.

Yet the field still operates far behind the capability frontier. Analyzing a single complex prompt can require hours of expert attention; methods that work on small models remain brittle at scale; and fundamental definitions remain imprecise. The gap between AI capabilities and AI understanding is widening, and regulatory pressure is mounting. The next five years will be decisive: either interpretability tools automate and scale to match frontier models, or the world will face unprecedented decisions about deploying systems whose inner workings remain genuinely mysterious — even to their creators.

14 of 14. References1.
  1. Understanding Mechanistic Interpretability in AI Models - IntuitionLabs - The January 2025 "Open Problems" paper further identified that core concepts like "feature" still la...

  2. Accelerating AI Interpretability - Federation of American Scientists - If AI systems are not always reliable and secure, this could inhibit their adoption, especially in h...

  3. Explainable AI: A Review of Machine Learning Interpretability Methods - Recent advances in artificial intelligence (AI) have led to its widespread industrial adoption, with...

  4. [2404.14082] Mechanistic Interpretability for AI Safety -- A Review - This review explores mechanistic interpretability: reverse-engineering the computational mechanisms ...

  5. Interpretability vs. explainability in AI and machine learning - Learn the key differences between interpretability and explainability in AI and machine learning, an...

  6. Interpretability versus Explainability in Deep Learning with ...

  7. Psychological Foundations of Explainability and Interpretability in Artificial Intelligence

  8. Investigating the Duality of Interpretability and Explainability ... - arXiv

  9. Explainability vs. Interpretability

  10. Attention Probes - Learn Mechanistic Interpretability - Probing classifiers test what information is encoded in a model's hidden states. The standard setup ...

  11. Interpretability Research - Anthropic - The mission of the Interpretability team is to discover and understand how large language models wor...

  12. Sparse Autoencoders Find Highly Interpretable Features in ... - One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanti...

  13. Scaling Monosemanticity - Paper summary

  14. On the Interpretability of Attention Networks - Attention mechanisms form a core component of several successful deep learning architectures, and ar...

  15. Tracing the thoughts of a large language model - Anthropic's latest interpretability research: a new microscope to understand Claude's internal mecha...

  16. Circuit Tracing: Revealing Computational Graphs in Language Models - We describe an approach to tracing the “step-by-step” computation involved when a model responds to ...

  17. On the Biology of a Large Language Model - We investigate the internal mechanisms used by Claude 3.5 Haiku — Anthropic's lightweight production...

  18. Demystifying AI: Understanding 'On the Biology of a Large ... - In a paper titled On the Biology of a Large Language Model, researchers from Anthropic explain how t...

  19. Anthropic’s latest paper, “On the Biology of a Large Language… | Anthony Butler - Anthropic’s latest paper, “On the Biology of a Large Language Model,” offers some of the clearest in...

  20. Open-sourcing circuit-tracing tools - Anthropic - Anthropic is an AI safety and research company that's working to build reliable, interpretable, and ...

  21. Part 2: 4. Interpretability (older version) - Neel Nanda discusses mechanistic interpretability and its possible applications to enabling alignmen...

  22. OpenAI Researchers Train Weight Sparse Transformers to Expose Interpretable Circuits - Weight sparse transformers expose small interpretable circuits, advancing mechanistic interpretabili...

  23. OpenAI found features in AI models that correspond to ... - TechCrunch - OpenAI researchers say they've discovered hidden features inside AI models that correspond to misali...

  24. Interpretability & Mechanistic Understanding - Chapter 10: Interpretability & Mechanistic Understanding. As LLMs become more capable, understanding...

  25. [PDF] Open Problems in Mechanistic Interpretability | Semantic Scholar - This forward-facing review discusses the current frontier of mechanistic interpretability and the op...

  26. Superposition, Memorization, and Double Descent - Anthropic - In a recent paper, we found that simple neural networks trained on toy tasks often exhibit a phenome...

  27. Polysemanticity and Capacity in Neural Networks - arXiv

  28. Linear representation hypothesis | AI Research Papers - The "linear representation hypothesis" in AI/ML research explores the idea that complex, non-linear ...

  29. How Anthropic Is Mapping Claude's Hidden Reasoning | Answer

  30. Anthropic Just Dropped a Stunning Paper — And It Might Change How We Understand AI | FrontPage - Anthropic just released one of the most fascinating and visually striking AI research papers of the ...

  31. Circuits Updates - September 2025

  32. Evaluating self-attention interpretability through human-grounded experimental protocol - Attention mechanisms have played a crucial role in the development of complex architectures such as ...

  33. Open Source Automated Interpretability for Sparse Autoencoder ... - Building and evaluating an open-source pipeline for auto-interpretability

  34. Automated Framework Enhances Neural Network Interpretability With Scalable Explanations - Researchers developed an automated system using large language models to interpret millions of featu...

  35. Aligning AI Through Internal Understanding: The Role of Interpretability

  36. Aligning AI Through Internal Understanding: The Role of ... - arXiv

  37. Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons - Large language models (LLMs) excel in various capabilities but pose safety risks such as generating ...

  38. [PDF] Expert Views on Aligning Explainable AI with the EU AI Act

  39. EU AI Act implementation: New obligations for general-purpose AI ... - The European Union’s Artificial Intelligence Act represents the world’s first comprehensive legal fr...

  40. Automated Approaches to Interpreting and Explaining Machine Learning Models - This paper proposes a novel framework for automating ML model interpretation and explainability acro...

  41. Transparency in AI Decision Making: A Survey of Explainable AI Methods and Applications - Artificial Intelligence (AI) systems have become pervasive in numerous facets of modern life, wieldi...

  42. Sparse autoencoders uncover biologically interpretable features in protein language model representations | PNAS - Foundation models in biology—particularly protein language models (PLMs)—have enabled ground-breakin...

Sections in this document