October 7, 2026:


When researchers try to understand how a large language model actually works — tracing which internal components are responsible for a given behavior — they rely on a technique called circuit discovery. The dominant approach, activation patching, works by surgically removing one component at a time and measuring what breaks. A team of five researchers posted a paper to arXiv this week that directly addresses a structural flaw in this approach, one that has quietly undermined interpretability research since 2023: the moment you remove a component to test it, the rest of the model routes around the damage and compensates, meaning you are never measuring the intact system — only its ablation-distorted shadow.
That compensation is known as the Hydra effect, named in a 2023 Google DeepMind paper by Thomas McGrath and colleagues. Like the mythological monster that regrows two heads for every one that is cut off, a transformer’s components respond to ablation by having other components increase their contribution to fill the gap. McGrath, who is now Chief Scientist at Goodfire — the AI interpretability startup that reached a $1.25 billion Series B valuation in February 2026 — documented that this compensatory behavior makes two seemingly obvious ways of measuring component importance far less correlated with each other than they should be.
The new paper — titled “Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models” and posted to arXiv on October 2, 2026, by Sankaran Vaidyanathan, Rafal Urbaniak, Emily Bunnapradist, Michelangelo Naim, and Daniel Waxman — introduces a formal solution drawn from a 20-year-old tradition in causal inference theory. The authors borrow the concept of a witness from Pearl and Halpern’s causal inference framework, where a witness is a set of variables whose values provide contextual information that resolves whether a given cause genuinely produced an effect. Applied to circuits, conditioning on witnesses lets researchers estimate a component’s genuine contribution without conflating it with the compensatory responses its removal triggers in other components.
The problem the paper addresses is best illustrated through GPT-2’s indirect object identification (IOI) circuit — the subject of a foundational 2022 IOI circuit paper by Wang and colleagues at Redwood Research and among the most thoroughly studied circuits in the field. Researchers discovered that GPT-2 uses a set of name-mover attention heads to copy the correct name to the output in sentences like “When John and Mary went to the store, John gave the bag to ___.” However, the IOI circuit also contains “backup name-mover heads” that are normally silent. When the primary name-mover heads are knocked out by ablation, the backup heads activate and take over. This is the Hydra effect in operation.
The consequence, as the new paper states, is that two natural measures of a component’s importance — unembedding-based and ablation-based — become far less correlated than researchers would naively expect. A circuit discovered by ablating components one at a time is, in part, a circuit of the ablated model, not the intact model. The backup mechanisms are not part of the model’s normal computation; they only emerge under the artificial conditions of the researcher’s intervention. Yet they are precisely what standard circuit-discovery tools end up measuring.
This is not limited to GPT-2 or the IOI task. Ablating a layer can trigger self-repair in subsequent layers, a pattern documented across multiple model families. The field’s major automated circuit-discovery tools — ACDC, Edge Attribution Patching (EAP), Edge Pruning, and others — all share the same foundational limitation: they score components in isolation, and “the joint effect of any component set includes cross-interactions for every combination of its members, which per-component scores cannot separate,” as the paper’s core argument states.
The paper’s theoretical contribution is the witness-integrated set effect (WISE): a family of causal estimands that take expectations over sets of causes and witnesses jointly, rather than evaluating components one by one. In the language of Pearl’s structural causal models, a witness is a set of variables W held at their actual values in the intact model, whose presence conditions the causal estimate so that the estimate reflects what component X actually does in the model’s normal operating context — not what it does when the model is already damaged by ablation.
This matters because the standard approach measures what is known in causal inference as the “total effect” of ablation — which includes both the component’s genuine contribution and the compensation it triggers everywhere else. WISE separates these by conditioning on witnesses: variables whose values pin down the state of the system in a way that neutralizes the confounding effect of compensatory responses. The approach is computationally tractable because expectations are taken over sets of components and witnesses rather than requiring enumeration of all possible interactions.
Building on WISE, the authors introduce JuntaLearner — a gradient-based circuit discovery algorithm named after the “junta” concept from combinatorics and social choice theory, which refers to a set of variables that jointly determine an outcome. The name is deliberate: relevant components need to be evaluated as a group, not in isolation. JuntaLearner iterates over varying-sized sets of components and witnesses, learning to rank components by their causal impact across the varying witness conditions. It is the first circuit-discovery algorithm designed around the combinatorial structure that the Hydra effect reveals.
Standard circuit evaluation in mechanistic interpretability relies primarily on faithfulness: the degree to which a discovered circuit reproduces the full model’s output when the rest of the model is ablated. The authors argue this single metric is insufficient — a circuit can score well on faithfulness by incorporating backup heads that only exist because the rest of the model was ablated, thereby building the artifact of measurement into the circuit itself.
The paper introduces a kill metric that measures necessity: how much damage removing the discovered circuit causes to the full model’s performance. A genuinely good circuit should be both sufficient (faithful) and necessary (lethal when removed). The authors also introduce measures of task specificity — whether the circuit is doing something task-relevant rather than implementing a generic model property — and assemble all of these into a composite circuit recognition score (CRS) that summarizes performance across circuit sizes while giving extra weight to compact circuits, on the principle that a genuine mechanistic explanation should be lean.
Evaluated across multiple tasks and models of increasing scale, JuntaLearner’s CRS results show higher mean scores than attribution patching baselines on all metrics. The advantage is most pronounced on faithfulness and kill measures — exactly the metrics where the Hydra effect’s distortion bites hardest.
Mechanistic interpretability is one of the most actively funded areas of AI safety research. MIT Technology Review’s 2026 breakthrough list named it among ten transformative technologies of the year. The White House AI Action Plan published in July 2025 explicitly called for interpretability investment as a strategic priority. Field-wide, roughly $75 million to $150 million per year in dedicated investment flowed into mechanistic interpretability globally in 2026, with Goodfire’s $150 million Series B alone representing one of the largest single-round commitments.
Anthropic — whose CEO Dario Amodei described mechanistic interpretability as “among the best bets to help us transform black-box neural networks into understandable, steerable systems” — used interpretability in Sonnet 4.5’s deployment, marking the first time interpretability research directly informed a production model release decision.
A circuit that is partly an artifact of ablation-induced compensatory behavior is not a real explanation of model behavior. It is a map drawn from a landscape that only exists when the map-maker is already intervening in it. Researchers using such circuits to study factual recall, reasoning, or potentially dangerous capabilities have been working from a systematically distorted picture of how the model actually operates. If those circuits inform safety interventions — model editing, activation steering, scalable oversight mechanisms — the interventions are designed against a ghost.
The paper tests JuntaLearner “across multiple tasks and models of increasing scale,” but it does not claim to solve the full scalability challenge facing mechanistic interpretability. Frontier models like GPT-4 or Claude Opus contain billions of parameters and hundreds of attention heads across dozens of layers, creating a combinatorial space that no circuit-discovery algorithm has yet managed to navigate comprehensively. The Hydra effect is one of several documented reasons why circuits discovered in toy-model settings do not cleanly transfer to production models at scale — others include superposition and polysemanticity, discussed in the Sharkey et al. open problems survey.
What WISE and JuntaLearner provide is a more principled foundation for the circuits that can be discovered, at whatever scale the computation allows. It is now possible, in principle, to test whether a discovered circuit survives interaction-aware scrutiny before drawing conclusions about model internals. That is a meaningful methodological advance even if the scalability frontier has not moved.
The 2026 International AI Safety Report, authored by over 100 experts from more than 30 countries and led by Turing Award winner Yoshua Bengio, explicitly identified interpretability as a critical enabler of safety while acknowledging that it cannot substitute for the broader infrastructure of evaluations, monitoring, and governance. WISE and JuntaLearner represent one piece of that infrastructure becoming more reliable.
The paper “Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models” is available at arXiv:2610.04017. Authors are Sankaran Vaidyanathan, Rafal Urbaniak, Emily Bunnapradist, Michelangelo Naim, and Daniel Waxman, with correspondence to sankaranv@cs.umass.edu and rfl.urbaniak@gmail.com.
The Hydra effect — named and documented in a 2023 Google DeepMind paper — describes the tendency of transformer language models to compensate when one component is removed or ablated during testing. When researchers knock out an attention head to measure its importance, other attention heads quietly ramp up their contributions to fill the gap. This means that standard circuit-discovery tools, which test one component at a time, are measuring how the model behaves under artificial damage rather than how it actually works in normal operation. Circuits built on these measurements may incorporate backup components that are never active when the model runs normally — making them artifacts of the measurement process rather than genuine explanations of model behavior. That matters for AI safety because researchers use these circuits to guide model editing, activation steering, and safety-evaluation frameworks: if the circuits are distorted, so are the interventions built on them.
Activation patching and its derivatives (ACDC, edge attribution patching, edge pruning) all score components in isolation — they measure one component at a time, then aggregate. The fundamental problem is that the effect of any set of components includes cross-interactions among all possible combinations of those components, which per-component scores cannot separate. WISE (witness-integrated set effect) takes expectations over sets of causes and witnesses jointly, borrowing a framework from Judea Pearl and Joseph Halpern’s actual causality literature in which a “witness” is a set of variables held at their actual values to condition the causal estimate. This lets WISE estimate a component’s genuine causal contribution in the intact model rather than in the artificially ablated model that standard methods create by their own intervention.
The Hydra effect has been a recognized limitation of circuit discovery since 2023, but the field largely treated it as a caveat rather than a structural threat. The practical consequence is that circuits discovered by isolation-based methods tend to be brittle: they look faithful when measured against the ablation condition that produced them, but their faithfulness degrades under more demanding evaluation because they are partly built from interactions that only exist in the distorted conditions of the experiment. This does not mean all prior circuit research is wrong — the IOI circuit in GPT-2, for example, was documented extensively enough that researchers identified backup heads as a distinct phenomenon rather than treating them as primary components. But it does mean that circuits used as the basis for safety interventions should be re-evaluated with interaction-aware methods before being treated as reliable causal maps. JuntaLearner provides a practical tool for doing that.
In combinatorics and social choice theory, a junta is a set of variables that jointly determine an outcome. The JuntaLearner algorithm is named deliberately to signal its core principle: the relevant components of a circuit must be evaluated as a group, not in isolation. This is not just terminological — it reflects the algorithmic design, which iterates over varying-sized sets of components and witnesses to learn which combinations jointly account for model behavior in a way that is stable under different witness conditions. The name is a compact statement of the paper’s critique of the field: if you have been evaluating components one at a time, you have been asking the wrong question.