Ablation Can Make Language Models Give the Wrong Answer About How They Work

A new arXiv paper argues that a common technique for mapping language models measures the wrong object: the ablated model, not the original one. The authors' answer is a framework called WISE — and a tool whose workings the coverage does…

Still life of an anatomical wire model of a head with one part removed, beside a rough clay copy examined instead — an illustration of ablation measuring the modified model rather than the original.
Illustration
Gift article

Ablation Can Make Language Models Give the Wrong Answer About How They Work

A new arXiv paper argues that a common technique for mapping language models measures the wrong object: the ablated model, not the original one. The authors' answer is a framework called WISE — and a tool whose workings the coverage does not explain.

The News

Five researchers — Sankaran Vaidyanathan, Rafal Urbaniak, Emily Bunnapradist, Michelangelo Naim and Daniel Waxman — uploaded the paper "Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models" to arXiv on October 2, 2026. The description here is based on TechTimes' coverage of the paper, published October 6, 2026 (TechTimes).

The claim is a serious one for the field: according to TechTimes' characterization, a widely used technique for understanding what happens inside language models has a structural flaw that has undermined interpretability research since 2023. The new contribution, according to the coverage, consists of a formalization of the problem — and a proposal for how to get around it.

The Mechanism: Why Ablation Can Measure the Wrong Thing

"Circuit discovery" means identifying which components of a model — neurons, attention heads, layers — carry out a particular task. The standard method is ablation: you "switch off" a component and see how much the model's performance drops. The larger the drop, the more important the component is assumed to be.

The problem is that models respond to ablation. When a component is removed, other components can step in and compensate. This phenomenon is called the Hydra effect, named in a 2023 Google DeepMind paper by Thomas McGrath and colleagues — after the monster that grew new heads when old ones were cut off. (TechTimes reports that McGrath is now Chief Scientist at the company Goodfire; this has not been verified against primary sources.)

The consequence: per-component scores conflate a component's real contribution with the intact model's compensation machinery. The measurement does not reflect what the component does in the original model, but what the system does when it is gone.

A Concrete Example: GPT-2 and the Backup Heads

The classic example is GPT-2 and the "indirect object identification" (IOI) task — sentences like "When John and Mary went to the store, John gave the backpack to ___" — mapped in a 2022 paper by Wang and colleagues. The IOI circuit contains "backup name-mover heads" that are normally silent. Only when the primary name-mover heads are knocked out by ablation do the backup heads activate and take over.

This means that an ablation study of the primary heads can underestimate their importance: the model already has a backup system that dampens the damage — and it is the backup heads, not the original function, that shape the measurement.

The Attack on Established Tools

According to TechTimes, the paper's authors argue that the major automated circuit discovery tools — ACDC, Edge Attribution Patching (EAP), Edge Pruning and others — share the same fundamental limitation: they score components in isolation. The core of the argument is rendered as follows: "the joint effect of any component set includes cross-interactions for every combination of its members, something per-component scores cannot distinguish" (TechTimes). In other words: interactions between components cannot be captured by measuring them one at a time.

The Proposal: WISE — and an Unexplained Tool

The paper's theoretical contribution is the witness-integrated set effect (WISE): a family of causal estimands that takes expectations over sets of causes and "witnesses" jointly, rather than evaluating components one by one. The "witnesses" are drawn from the causal inference framework of Judea Pearl and Joseph Halpern, and are meant to distinguish genuine contributions from compensatory responses.

The coverage also mentions a tool called JuntaLearner, presented alongside WISE as part of the new contribution. How the tool works does not emerge from the available coverage. Nor are any empirical results or benchmark figures from the paper visible in the available reporting.

What Remains Unresolved

Several caveats should accompany this story:

  • One secondary source. All technical claims here rest on TechTimes' rendering. The underlying arXiv paper should be read before technical details are treated as fully documented.
  • The authors' argument, not an established effect. That WISE and JuntaLearner "fix" circuit discovery is the authors' claim. The methods have not been validated by independent empirical work, and no results are available in the available coverage.
  • Career detail unverified. The information that McGrath leads research at Goodfire comes from the same secondary coverage.

The story is nonetheless anchored in established research: the Hydra effect has been a known phenomenon since the 2023 DeepMind paper and the 2022 Wang paper, and the critique targets well-established tools such as ACDC, EAP and Edge Pruning. If WISE holds up in practice, it could change how circuits in language models are mapped — but that must be confirmed by independent empirical work.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.