TaH2
Better test-time scaling
for looped models.
Adaptive depth improves
test-time scaling.
With more test-time compute and a higher depth ceiling, TaH2 keeps improving where fixed-depth baselines flatten out.
Extra iterations help
only some tokens.
Fixed-depth looping spends compute on every token—even when another pass barely helps, or makes the prediction worse.
The token-level analysis ↗Little change: |loss reduction| ≤ 10⁻³. Fixed-depth models post-trained from Qwen3-1.7B-Base.
The question is not just “Is this token hard?”
It is “Will another iteration help?”
Learning when to iterate.
The backbone and decider learn together, using the actual gain from another iteration.
Keep the input in the loop.
A learned updater combines the original embedding with the previous hidden state.
Learn from iteration gains.
Online, cost-sensitive depth labels teach the decider where extra compute helps.
Use the predictions along the way.
Stopping probabilities weight predictions across the executed iterations.
What changes from TaH to TaH2?
TaH uses staged training and offline top-1 mismatch labels. TaH2 jointly trains the backbone, updater and decider with lookahead depth supervision: labels track gains on the current backbone. Both training and inference follow the decider’s thresholded decisions. Extended duo-causal attention preserves causality across token positions and iteration depths.
At inference, a token stops when its continuation probability is below 0.5, or when it reaches the depth ceiling. Training-only lookahead is omitted. The updater and decider add fewer than 3% parameters at the studied scales.
Adaptive depth, token by token.
A live generation trace shows predictions across iterations, with committed tokens colored by their final depth.
Compute, accuracy,
and latency.
Steeper sequential and parallel test-time scaling, lower loss throughout training, real wall-clock gains in serving, and depth that follows the content.
Ratios relative to Standard.
Gains across tasks
and model scales.
From mathematical reasoning to code, knowledge and tool use.
TaH2 M=2 · Separate training setup from the 1.7B study.
All nine benchmarks · 4B and 8B
Accuracy across nine benchmarks. Labels show TaH2’s gain over Standard in percentage points.
Team
Meet the researchers.
*Equal contribution †Corresponding authors
Citation
Cite our work.
If you find TaH2 useful for your research, please consider citing our paper.
@misc{you2026tah2,
title={Improving Test-Time Scaling with Adaptive Looped Transformers},
author={You, Yichen and Fu, Tianyu and Feng, Aosong and Lv, Xingtai and Ning, Xuefei and Ding, Ning and Wang, Yu},
year={2026}
}