MIT and Sakana AI: SIFT reached 35.1 percent on Polyglot after only 30 steps

Researchers at MIT and Sakana AI describe a framework that lets a language model judge candidates instead of running expensive benchmark tests on every change — with reported results surpassing Darwin Gödel Machine at around one-tenth of…

Illustration: dozens of identical aluminum blocks on a dark table, a single sculpted stamp pressing down on one selected block marked with an acid-yellow accent — a metaphor for a judge selecting candidates instead of testing each one.
Illustration
Gift article

MIT and Sakana AI: SIFT reached 35.1 percent on Polyglot after only 30 steps

Researchers at MIT and Sakana AI describe a framework that lets a language model judge candidates instead of running expensive benchmark tests on every change — with reported results surpassing Darwin Gödel Machine at around one-tenth of the resources.

Researchers at MIT and the Japanese lab Sakana AI have described a framework called SIFT, short for Self-Improvement via Fast Tree-search, that is intended to make it substantially cheaper to let coding agents improve themselves. According to a report from Cryptobriefing (Diego Almada Lopez, October 2, 2026), SIFT achieved 35.1 percent on the coding benchmark Polyglot after only 30 expansion steps. The comparison point is Darwin Gödel Machine (DGM), an earlier approach that reached 30.7 percent on the same benchmark — but only after 80 nodes.

All figures below are as reported from the article, not independently verified.

Why evaluation is the bottleneck

Self-improving coding agents work by generating many candidate changes to their own code and selecting the best ones. The problem is selection: determining whether a change actually makes the agent better traditionally requires a full benchmark run for each candidate. That makes the search expensive in compute and time, and limits how broadly one can search.

How SIFT works

SIFT replaces the expensive runs with a judge. Instead of benchmarking every change, the framework has a large language model compare two candidate modifications and say which one looks better, according to Cryptobriefing's description of the article.

The judge's verdicts are then combined with a regularized Bradley-Terry model, a statistical method for turning pairwise preferences into an overall ranking. The actual evaluation runs asynchronously, so that only the candidates that make the shortlist face the real benchmark.

The point, in other words, is not to eliminate benchmark runs entirely, but to use them far more selectively: the language model sifts, the benchmark confirms.

The reported numbers

In addition to the Polyglot result, improvements are reported on two other benchmarks:

  • TerminalBench 2.1: from 29.2 to 36.7 percent, an increase of 7.5 percentage points.
  • SWE-60: from 40.0 to 52.1 percent, an increase of 12.1 percentage points from the starting point.

The cost side may be the most striking. A SIFT configuration built on the open model Qwen3-Coder-30B reportedly completed the entire search in 224 CPU hours with roughly $34 in API costs — around one-tenth of the resources DGM consumes, according to the report. The article does not state DGM's exact resource use as a detailed basis for that comparison.

Who is behind it, and when

The article is authored by Xinghong Fu at MIT, together with Aravinth Kulanthaivelu and Yutaro Yamada, with Sakana AI as the collaborating lab. It was posted to arXiv around September 18, 2026, and coverage spread in the weeks following.

What it means

If the numbers hold, SIFT changes the economics of research on self-improving agents. Searches that previously required large compute resources become accessible to smaller teams — the reported configuration also uses an open-weights model, not a closed commercial service. For labs experimenting with agent architectures, cheaper evaluation means more searches, broader searches, or both.

It is worth emphasizing that this is a paper result, not documented production use.

The main objection: the judge

The central unresolved risk, which Cryptobriefing also points to, is judge quality. If the language model prefers the wrong patches, the cheaper search will be steered wrong — faster, but in the wrong direction. How often this happens, and how well the Bradley-Terry ranking counteracts it, is not quantified in the available material.

Open questions

Several details remain unconfirmed: the exact arXiv identifier, the article's full title, and the complete author list have not been verified. DGM's exact resource basis behind the "one-tenth" comparison is not specified in the report, and the error rates of the LLM judge are unknown. Since all figures come from a single secondary source citing the article, independent verification remains outstanding — particularly on whether the results generalize beyond the three benchmarks reported.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.