Small LIFT model beats same-size transformers trained on eight times more data — authors say

A new research paper proposes a way to address one of the transformer architecture's most fundamental limitations: that information between generation steps can only flow through the decoded token.

Illustration: dozens of colored threads squeezed through a tiny brass ring and exiting as a single braided cord, a metaphor for how information in a transformer can only flow through one channel between generation steps.
Illustration
Gift article

Small LIFT model beats same-size transformers trained on eight times more data — authors say

A new research paper proposes a way to address one of the transformer architecture's most fundamental limitations: that information between generation steps can only flow through the decoded token. The work, published as a preprint on arXiv, presents results from models ranging from 135 million to 1 billion parameters — but everything so far rests on the researchers' own reported numbers.

What is new

On 29 September 2026, at 17:57:40 UTC, Dor Tirosh, Ido Amos and Mor Geva uploaded version 1 of the preprint "Pretraining Latent Information Feedback Transformers with Teacher Supervision" to arXiv (arXiv:2609.38149). The paper introduces LIFT — the Latent Information Feedback Transformer — an architecture and training method that, according to the authors, allows language models to propagate state across generation.

In practice, this means the model is no longer dependent on squeezing everything it has learned between one generation step and the next into a single output token. Instead, it passes a learned latent state directly from one step to the next — something otherwise known from recurrent networks, but here in a way that preserves the transformer's parallel pretraining.

The problem LIFT targets

In a standard transformer, information flows in one direction: representations in the deeper layers are never passed back to the shallower layers, and the only channel for information to flow across generation steps is the decoded token. That is how the authors describe it in the abstract. This narrow channel, they argue, forces models to recompute intermediate results and to discard alternative continuations.

In other words: when the model generates a word, everything else it "thought" along the way is lost, unless it happens to find its way into the chosen token. LIFT, according to the authors, aims to remove this bottleneck during pretraining.

How the mechanism works

The core trick in LIFT is to turn learning of recurrent state into a prediction problem with teacher forcing. Each input token is paired with an information-dense state derived from the next-token distribution of an off-the-shelf, fully trained language model. The model is then trained to predict both the next token and the next state.

This yields a practical advantage: because the input states are precomputed, pretraining remains fully parallel across positions — the enormous efficiency gain that makes transformers trainable at scale is not lost.

At inference, the picture changes. The model no longer receives teacher states; instead, its own predicted states are fed back into the next step. This incurs an increase in computational cost — but, according to the authors, a modest one, and the overhead decreases with model size. It is a trade-off, then: parallel training with teacher states, feedback of self-generated states at generation time.

What the researchers report

All of the results below come from the authors' own abstract and are not independently verified.

According to the authors, experiments with pretrained models from 135 million to 1 billion parameters show that LIFT consistently outperforms standard transformers and baselines on three types of tasks under token-matched budgets: language modeling, downstream reasoning tasks, and procedural tasks. In comparisons with compute-matched transformers, LIFT is, by their account, on par with or ahead.

The most detailed result concerns a controlled study of a state-tracking task — a task type in which the model must keep track of an internal state across a sequence. There, the authors report that a small LIFT model outperformed same-size transformers that had been trained on eight times more data. More striking still: this held even when LIFT was trained with teacher states drawn from a transformer that itself failed the task. In other words, the method can reportedly extract useful state information from a teacher that fails at the task itself — suggesting that the teacher's next-token distribution contains more structure than its final answers reveal.

What is not yet shown

Several substantial questions remain open, and they should weigh heavily in how the results are read:

  • Scaling. The evidence covers models up to 1 billion parameters. Whether feeding back self-predicted states still works at far larger scales — or whether small errors in the self-generated states accumulate over long generations — is unresolved in the available material.
  • Quality of self-generated states at inference. During training, the model sees accurate teacher states; during generation, it sees its own predictions. The abstract summarizes this transition but does not quantify how good the states actually are.
  • Practical cost. The mechanism requires teacher states to be precomputed from an existing pretrained model for the entire training dataset. The abstract describes the inference overhead but not the practical cost of the precomputation itself.
  • Independent verification. Detailed benchmark tables, baselines, model configurations and any code availability cannot be verified from the abstract alone, and there is so far no secondary coverage or outside scholarly commentary on the work.

Why the trade-off may matter

If the results hold up under verification and scaling, LIFT points to an interesting middle position in the debate over language model architectures. The transformer's strength is full parallelism during training; the recurrent model's strength is explicit state across steps. LIFT proposes to get both: state propagation as a prediction task during training, feedback as a small cost at run time.

The open challenge lies in the transition between the two regimes. The models are trained with teacher states but run on their own — and whether that gap closes cleanly at larger scale is the question that must be answered before anything else in the results can be read as more than a promising signal from a fresh piece of research.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.