Skip to content

The opportunity

One response.
Different models.

Token-level routing lets small and large models collaborate within a single response. The serving engine must keep up with every switch.

Where does routing happen?

Model assignment at three granularities

Turn 1 · Token 1

Switch models within a response, at the granularity of a token or a span.

Small modelLarge model
Illustrative model assignments for three turns of one session; not a measured generation trace.

Independent progress.

A fast model should keep decoding while a slower peer handles routed tokens.

Useful batches.

Irregular token arrivals need a scheduler that brings requests together at the right time.

A simple interface.

The routing policy describes one request. The runtime manages batching, handoff and cache state.

The engine

Each model moves
at its own pace.

Request-centric programming.
Model-centric execution.

One serving endpointRequest admission & streaming
S

Small model

Entry subserver

SchedulerLLM runnerRouter
Decoding
ABC
Step 1
Local KV cacheRequest A · runningA
send / receiveToken suffix
+ routing state
L

Large model

Peer subserver

SchedulerLLM runnerRouter
Decoding
Idle
—
Local KV cacheIndependent model state

Decode a local batch.

The small model advances requests A, B and C in one batch.

Decode
Illustrative handoff; animation speed does not represent model latency. Each subserver has its own decoding loop. Only token suffixes and routing state travel between models; KV caches stay local. The architecture extends to more than two models.
Client–server loop

Admit requests and stream completed tokens.

Decoding loop

Schedule local batches and execute model steps.

Inter-model loop

Send, receive and resume routed requests.

Resume without starting over.

A pending request keeps its serving state and KV slot. Returning tokens are appended without repeating prefix matching or KV allocation.

RunningPendingRunningor finished

Programming interface

Three functions.
One routing policy.

Describe when a request changes models, what travels with it, and how decoding resumes.

route(batch, result)

Choose the next model.

After each forward pass, return one model name per request. Returning the current model name continues local decoding. Existing two-model schedulers also accept Boolean decisions.

R-Stitch example: True delegates to the only peer; False keeps decoding locally.

R-Stitch excerpts from the released implementation. Configuration and surrounding class definitions are omitted.

Five evaluated policies, the same interface
PolicyModelsRouting decision
CITER2Low confidence routes to a peer for one token.
R2R2A learned router predicts when the models would diverge.
R-Stitch2Entropy controls switching in both directions.
Co-LLM2A learned deferral signal requests a peer token.
ME2+Models are selected token by token using ensemble weights.

The released code also includes GlimpRouter, query-level routing and a random baseline. See the supported schedulers.

Quick start

Install from source and launch an R2R server with the Qwen3 0.6B / 32B configuration.

git clone https://github.com/thu-nics/TokenRouter.git
cd TokenRouter
conda create -n tokenrouter python=3.10
conda activate tokenrouter
pip install -e .
python -m tokenrouter.launch_server \
  --config_path config/r2r/Qwen3-0.6B_Qwen3-32B.yaml

Once the server is running, use the OpenAI client example or the requests example. The Python engine example runs directly without a separate HTTP server.

The README covers multi-node serving; the reproduction guide includes configurations, workloads and benchmark commands.

Delayed batching

A small wait.
A better batch.

Starting immediately can leave the next arrival waiting for a whole decoding step. A short, controlled wait brings routed requests into the same batch.

Same tokens. Different schedules.

Three requests · five tokens per request

Scheduling timelineA conceptual comparison of four execution strategies.
t = 0.0
Small model

Ready

Large model

Idle

Completed

0 of 3 requests

—Mean completion time
Small model · 1 unitLarge model · 3.5 units
Illustrative schedule, not a benchmark
Fixed routing sequences and latencies from the paper’s scheduling example. Each block is one token; shared start times indicate a batch. Delayed batching uses Bsmall = 1 and Blarge = 2.

Measured threshold sweep

The threshold matters.

A larger batch is useful until waiting leaves too little work in flight. The best threshold depends on concurrency and routing behavior.

Throughput · output tokens/s

R2R · Qwen3-0.6B / 32B · AIME2024. Measured values from the supplied threshold sweep.
How is the threshold selected?

A discrete-time Markov chain models queued requests, active batches and remaining execution time for each subserver. Given concurrency, routing probabilities and model step latencies, TokenRouter searches the feasible thresholds for the highest expected throughput.

B* = arg maxB   E[committed tokens] / E[elapsed time]

The condition ∑i(Bi − 1) < N avoids a state where all requests wait in queues. The analytical model assumes stationary routing probabilities and fixed per-step model latencies; it does not model every nonstationary routing pattern.

Serving efficiency

Routing gains.
Delivered in tokens per second.

Five algorithms, three workloads, and two baselines. TokenRouter accelerates the system that executes the routing policy.

Concurrency
Metric

High-effort reasoning

Concurrency 4 · output tokens/s · higher is better

Official CodeStd. ServingTokenRouter
Ratios compare TokenRouter with the stronger available baseline for each algorithm. All bars share a linear scale starting at zero.

Exact values for this selection
Models, workloads and measurement setup

Models & hardware

CITER, R2R, R-Stitch and Co-LLM use Qwen3-0.6B / 32B. ME adds Qwen3-8B, with weights 0.5 / 0.3 / 0.2.

Experiments use an 8 × A100-80GB host. Two-model runs share two GPUs via CUDA MPS: the small model uses TP1 and the large model TP2. ME places each model on one GPU.

Workloads

Low-effort: AIME2024, ~100-token inputs and a 2,048-token output cap.

High-effort: selected AIME2024 problems, ~100-token inputs and an 8,192-token output cap.

Agentic: multi-turn SWE-Smith trajectories, ~8,192 input tokens and a 1,024-token output cap.

Baselines & metrics

Official Code: the algorithm’s released implementation. Unavailable entries are shown as N/A.

Std. Serving: one SGLang server per model, returning one token per call to an external dispatcher.

Throughput is total generated output tokens divided by elapsed time. Latency is end-to-end request completion time. TokenRouter builds on SGLang 0.5.1.

Scaling concurrency

More throughput.
Responsive generation.

As concurrency grows, asynchronous execution turns more in-flight requests into useful model work.

Generation speed vs. total throughput

R2R · concurrency 1–16

TokenRouterR2R official
Throughput and per-user generation speedTokenRouter and the official R2R implementation across concurrency levels one through sixteen.
8
Each point is a measured concurrency setting; lines connect settings in order. Speed is per-user token generation rate. Throughput aggregates all requests.
18.58×

Throughput with comparable per-user speed.

TokenRouter at concurrency 16 reaches 630.57 tokens/s and 39.41 tokens/s per user. R2R official at concurrency 1 reaches 33.94 tokens/s and 34.87 tokens/s per user. This comparison uses different concurrency levels.

Inside the gains

The runtime matters
at every step.

CUDA graph support, asynchronous execution and delayed batching build on one another.

Cumulative system optimizations

R2R · concurrency 8

Output tokens/s
Each row includes the preceding optimizations. The official R2R implementation reaches 134.90 tokens/s; the complete engine reaches 372.48 tokens/s (2.76×). These are sequential ablations, not independent additive effects.

Extend CUDA graphs.

Capture a wider range of extend lengths encountered when a routed request resumes.

Keep routing local.

Run the router alongside the model to avoid a separate IPC round trip on every token.

Preserve serving state.

Retain pending requests and coordinate the KV budget across colocated models.

Different small / large model pairs
R2R · throughput (output tokens/s) across Qwen3 model pairs
Small / large modelR2R officialTokenRouterGain
0.6B / 8B105.10336.983.21×
0.6B / 32B77.08197.402.56×
1.7B / 8B103.37285.412.76×
4B / 8B106.67211.811.99×
Across nodes, with the same model resources

Qwen3-0.6B uses one GPU and Qwen3-32B uses two GPUs (TP2), without GPU overlap. Models are deployed on one or two nodes connected by 3.7 GB/s RoCE. The following measurements use low-effort reasoning.

Throughput in output tokens/s · independent deployment experiment
ConcurrencyDeploymentR-StitchR2RCITER
1Single node47.5374.90108.65
Two nodes46.1969.9293.99
4Single node116.14228.94308.36
Two nodes104.20202.43277.56
Algorithms in their original model and task settings
Concurrency 4 · latency in seconds · throughput in output tokens/s
Algorithm / original settingImplementationThroughputTTFTLatency
R2RDeepSeek-R1-Distill-Qwen 1.5B / 32B
AIME
LLM-only145.300.13411.98
Official Code89.620.11751.15
TokenRouter244.560.11270.19
CITERQwen2 1.5B / 72B
CommonsenseQA
LLM-only123.720.0832.68
Official Code17.160.0367.64
TokenRouter149.310.0360.48
Co-LLMTuned LLaMA2 7B / 70B
GSM8K
LLM-only134.790.06913.39
Official Code3.461.14247.61
TokenRouter76.020.06711.46
R-StitchL1-1.5B-short / QwQ-32B
AIME · official code unavailable
LLM-only150.410.15360.76
TokenRouter140.580.067151.30

Team

Meet the researchers.

*Equal contribution†Corresponding author

Citation

Cite our work.

BibTeX
@article{fu2026tokenrouter,
  title={TokenRouter: Efficient Serving System for Token-Level LLM Routing},
  author={Tianyu Fu and Tengxuan Liu and Ruoxi Wang and Yixin Dong and Yi Ge and Yichen You and Yu Wang},
  journal={arXiv preprint arXiv:2610.12242},
  year={2026},
}