TaH2

Better test-time scaling
for looped models.

THE INSIGHT

Adaptive depth improves
test-time scaling.

With more test-time compute and a higher depth ceiling, TaH2 keeps improving where fixed-depth baselines flatten out.

Qwen3-1.7B · AIME24–26 · avg@32
Standard + baselines · 4K–16KAIME24–26 · 1.7B · avg@3216K
Accuracy (%)
a

Better test-time scaling

4K–16K · fitted scaling trends

+53% slope
Test-time scalingMean AIME24–26 accuracy at output cutoffs from 4K to 16K. Fitted accuracy–log-compute slope: Standard 1.79, TaH2 M=2 2.74 points per compute doubling.
b

More depth, more headroom

16K output cutoff · varying depth ceilings

+3.9 pts
Iteration-depth scalingAt a 16K output cutoff, Standard reaches 10.97%. TaH2 reaches 13.78%, 14.13% and 14.86% at ceilings 2, 4 and 8, using 1.35, 1.82 and 3.31 times Standard decoding compute.

TaH2 turns more compute and greater depth into higher accuracy.

The observation

Extra iterations help
only some tokens.

Fixed-depth looping spends compute on every token—even when another pass barely helps, or makes the prediction worse.

The token-level analysis ↗
What changes after more iterations?
Iterations 1 → 23.42M validation tokens

Little change: |loss reduction| ≤ 10⁻³. Fixed-depth models post-trained from Qwen3-1.7B-Base.

The question is not just “Is this token hard?”
It is “Will another iteration help?”

THE METHOD

Learning when to iterate.

The backbone and decider learn together, using the actual gain from another iteration.

Explore the method ↗
ARCHITECTURELeft: token-level depth · Right: one training step
TaH2 architecture. Left: four tokens execute different numbers of iterations; the decider continues or stops each token, and stopped tokens take one lookahead iteration during training. Right: at each iteration the shared LLM backbone produces a hidden state; a head predicts the next token and a decider predicts the continue probability. Outputs across executed iterations are mixed with stopping weights and trained with a cross-entropy loss, while the decider is trained with a BCE loss against the lookahead gain.
A learned decider gives each token the depth it needs, supervised by the gain from one more iteration.
01 / REINJECT

Keep the input in the loop.

A learned updater combines the original embedding with the previous hidden state.

02 / SUPERVISE

Learn from iteration gains.

Online, cost-sensitive depth labels teach the decider where extra compute helps.

03 / COMBINE

Use the predictions along the way.

Stopping probabilities weight predictions across the executed iterations.

What changes from TaH to TaH2?

TaH uses staged training and offline top-1 mismatch labels. TaH2 jointly trains the backbone, updater and decider with lookahead depth supervision: labels track gains on the current backbone. Both training and inference follow the decider’s thresholded decisions. Extended duo-causal attention preserves causality across token positions and iteration depths.

At inference, a token stops when its continuation probability is below 0.5, or when it reaches the depth ceiling. Training-only lookahead is omitted. The updater and decider add fewer than 3% parameters at the studied scales.

Demo

Adaptive depth, token by token.

A live generation trace shows predictions across iterations, with committed tokens colored by their final depth.

Each token can take a different number of iterations. Colors indicate the final iteration depth.
THE EVIDENCE

Compute, accuracy,
and latency.

Steeper sequential and parallel test-time scaling, lower loss throughout training, real wall-clock gains in serving, and depth that follows the content.

Qwen3-1.7B · AIME24–26 · avg@32
Test-time scaling & training loss
TEST-TIME SCALINGOutput cutoff 4K
Accuracy (%)
TaH2 scales better with decoding compute. Lighter segments extend beyond the 16K training length.
TRAININGTraining step 0
SFT validation loss
TaH2 maintains lower validation loss throughout training. Insets show the final gap to Standard.
REAL-WORLD SERVINGMini-SGLang · A800 · AIME26
Runtime efficiencyBatch 1

Ratios relative to Standard.

Accuracy vs. measured end-to-end latency
Accuracy vs. latency
TaH2 achieves better scaling and higher peak accuracy in real-world serving. Lighter segments extend beyond the 16K training length.
TOKEN-LEVEL DEPTHTaH2 M=8 · three sampled correct responses
Text and depth appear together; darker colors mean more iterations.
BEYOND ONE SETTING

Gains across tasks
and model scales.

From mathematical reasoning to code, knowledge and tool use.

+3.2 pts average across nine benchmarks

TaH2 M=2 · Separate training setup from the 1.7B study.

All nine benchmarks · 4B and 8B
Accuracy of Standard and TaH2 at 4B and 8B on nine benchmarks

Accuracy across nine benchmarks. Labels show TaH2’s gain over Standard in percentage points.

Team

Meet the researchers.

1Tsinghua University    2Yale University
*Equal contribution    †Corresponding authors

Citation

Cite our work.

If you find TaH2 useful for your research, please consider citing our paper.

BibTeX
@misc{you2026tah2,
  title={Improving Test-Time Scaling with Adaptive Looped Transformers},
  author={You, Yichen and Fu, Tianyu and Feng, Aosong and Lv, Xingtai and Ning, Xuefei and Ding, Ning and Wang, Yu},
  year={2026}
}