10 October 2026 · By Hadi Ataei
TokenRouter: Making Token-Level Model Routing Fast

Most AI services pick one model for each request: a cheap small model for easy questions, an expensive large one for hard questions. But difficulty varies inside a single answer too. Most tokens in a long reasoning chain are easy ("the", "=", a repeated number), and only a few are hard decisions. Token-level routing exploits this by letting a small model write most tokens and calling in a large model only for the difficult ones. A paper submitted to arXiv on 8 October 2026, "TokenRouter: Efficient Serving System for Token-Level LLM Routing", is not about a new routing rule. It is about why token-level routing is slow in practice, and how to fix that. It was among the top papers on Hugging Face's daily list on 9 October, and the arXiv listing says it was accepted to NeurIPS 2026.
Background: request-level versus token-level routing
Routing algorithms decide which model handles what. At the request level, a router reads the prompt and sends the whole job to one model. At the token level, the choice is made again after every generated token, using signals such as how confident the small model is. The paper evaluates five published token-level methods, for example CITER (send low-confidence tokens to the larger model), R2R (send tokens predicted to diverge from the large model), R-Stitch (switch by entropy) and Co-LLM (a learned deferral). In the example the authors cite, R2R sends only about 5% of tokens to a 32B-parameter model, and the rest to a model of about 1.5B, while matching the large model's quality.
The problem: the math is cheap, the serving is not
On paper, this saves a lot of computation. In practice, the authors explain, standard LLM serving systems are built to run one model per request. Switching models after nearly every token causes two kinds of waste:
- Synchronization overhead: the models keep waiting for each other as control passes back and forth.
- Batch admission delay: GPUs are efficient when they process many requests in a batch. If requests keep hopping between two models, each model sees a thin trickle of arrivals and rarely gets full batches.
So the headline saving from using a small model is largely eaten up by the serving cost.
The design: request-centric programming, model-centric execution
TokenRouter separates how you write a routing method from how it runs.
- For the developer: you describe routing from the point of view of a single request with three functions.
route()decides where to go after each step (from logits or hidden states),send()packages the request state and the tokens the other model has not seen yet, andreceive()converts the reply back. - For the system: each model runs as its own autonomous subserver with decoupled loops for client communication, local decoding, and hand-offs to the other model. Requests are dispatched asynchronously.
- Cache handling: a "pending" state keeps a request's key-value (KV) cache while it is on the other model, which avoids recomputing the prefix on every switch. GPU memory is shared between the cooperating models automatically, based on their weight footprints.
- Delayed batching: the scheduler deliberately waits until enough requests have accumulated before launching a batch, trading a small queue wait for much better batch efficiency. The authors choose the threshold with a mathematical model of throughput built on a discrete-time Markov chain.
What the results say
The paper tests Qwen3 model pairs (0.6B with 32B as the main pair, plus 1.7B with 8B, and 4B with 8B) on workloads including AIME 2024 maths problems with long outputs and SWE-Bench-style agent traces. As reported by the authors:
- Throughput: 2.01 to 64.15 times higher decoding throughput than existing systems, across routing algorithms, workloads and model pairs.
- Examples at concurrency 4: R2R at about 245 tokens per second against about 90 for its official implementation (2.73 times); CITER at about 149 against 17 (8.7 times); Co-LLM at about 76 against 3.5 (about 22 times).
- Latency: end-to-end latency 2.03 to 63.64 times lower, with time to first token comparable to the baselines.
- Where it comes from (R2R, concurrency 8): CUDA graphs give 1.71 times, asynchronous execution adds a further 1.29 times, and delayed batching a further 1.25 times, for 2.76 times combined.
Note that these gains are in serving speed for the same routing rule. TokenRouter does not make the routing decisions better; it makes them fast enough to be worth using.
Caveats
- The large ratios, such as 64 times, come from comparing against the original research implementations of the routing methods, which were not designed as serving systems. Against a well-tuned production stack the gap would likely be smaller, and the authors' smaller 2 to 3 times figures are probably the fairer headline for R2R.
- The throughput model assumes the number of tokens between two switches follows a geometric distribution. The authors say this holds for most algorithms and leave other cases for future work.
- Experiments use Qwen3 models on a handful of workloads, and quality is determined by the underlying routing method, not by this paper.
Why it matters if you are learning AI
It is a good reminder that an idea that looks efficient on paper can fail on real hardware. Understanding batching, KV caches and scheduling is as important for AI engineering as understanding the models themselves. Routing also connects to ideas we have covered recently: the mixture-of-experts design behind Mistral Large 4 routes tokens to experts inside one model, while TokenRouter routes them between separate models. And, like ALoDLM, it is about spending computation unevenly where it is needed. If you want the foundations, see What Is a Transformer? and our Introduction to Natural Language Processing course.
What to watch next
- Comparisons against optimized production serving stacks, not only the original research code.
- Whether commercial inference providers adopt token-level routing.
- Routing between more than two models and across different model families.
Sources: arXiv:2610.12242 (submitted 8 October 2026, cs.CL) and its full-text version. The authors list code at github.com/thu-nics/TokenRouter. This is an independent explainer, not written by the paper's authors.



