Qatar AI Institute
← Back to articles

7 October 2026 · By

Distillation Across Tokenizers: Why Less Supervision Can Work Better

Distillation Across Tokenizers: Why Less Supervision Can Work Better

Training a small, cheap "student" model to imitate a large "teacher" is called knowledge distillation, and it is one of the main ways strong models are shrunk into practical ones. A paper submitted to arXiv on 6 October 2026, "Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability", asks a surprisingly practical question: when teacher and student read text differently, is it better to learn from everything you can match, or only from the parts you can match reliably? Its answer, as the title hints, is the second. It was the top paper on Hugging Face's daily list on 7 October.

Background: distillation, on-policy, and tokenizers

Three ideas are needed.

  • Distillation. The student is trained to make its predicted probability distribution over the next token match the teacher's, rather than just learning the correct answer.
  • On-policy. Instead of copying the teacher's own sentences, the student generates its own answers and the teacher grades each token it chose. In this paper the training objective is a reverse KL divergence, a measure of how far the student's distribution is from the teacher's, evaluated on the student's own generations. Because the student learns from the situations it actually gets into, this tends to work better in practice than offline copying.
  • Tokenizers. A tokenizer chops text into pieces (tokens) from a fixed vocabulary. Different model families chop differently. (Our article What Is BERT? shows how one tokenizer works.)

The problem: teacher and student speak different token languages

If teacher and student share a tokenizer, comparing their predictions is easy. If they do not, two things go wrong. The same text is split at different places, and the two vocabularies contain different entries. The authors measure the vocabulary overlap between three teacher-student pairs at only about 39% to 65%. So how do you compare "what the teacher thinks comes next" with "what the student thinks comes next"?

Two rows of tokens for the same sentence, one per model. Where both models use exactly one token for the same piece of text, a line joins them (strict positions). Where one side splits a word into several tokens, the tokens are marked as mismatch spans. Three boxes summarize the findings: strict positions alone cover 85.6 to 97.0 percent of student tokens, top-16 shared tokens keep at least 96 percent of the gain, and adding mismatch spans lowered accuracy.
Strict positions versus mismatch spans. The sentence and splits are invented for illustration; the percentages come from the paper.

Two strategies

Earlier methods try to recover as much supervision as possible, including from the awkward places where one model uses several tokens for what the other writes as one (mismatch spans). The authors argue for the opposite. Compare only at strict positions: places where each side has exactly one token covering exactly the same stretch of text. They then compare the two distributions only over tokens that both vocabularies contain, after renormalizing.

That sounds like throwing information away, but the authors show it throws away little:

  • Strict positions cover 85.6% to 97.0% of the tokens the student generates, depending on the model pair.
  • At those positions, 99.7% to 99.9% of the teacher's probability falls on tokens in the shared vocabulary, and 99.0% to 99.8% for the student. In other words, the shared vocabulary holds almost everything that matters.

Even smaller: the top 16

The authors go one step further. At each position they use only the 16 shared tokens the student currently rates highest, and compute the reverse KL on just those. They report that this keeps at least 96% of the improvement that using the full shared vocabulary gives over the undistilled student, at lower cost.

The surprise: more supervision hurts

When the authors add a loss on the mismatch spans (a mean-squared-error term on span log-probabilities, which brings supervision coverage to 100%), accuracy goes down for all three model pairs, by roughly 0.3 to 1.2 percentage points. They explain why with a diagnostic: they compare the direction of the gradient from the mismatch-span loss with the direction from the strict loss using cosine similarity. The values are close to zero, and sometimes negative, so the extra signal does not point the same way as the reliable signal. A control, splitting the strict positions into two random halves, gives a cosine similarity above 0.2, so the disagreement is not just noise. They also find the extra gradient grows relative to the strict one as training goes on, so the weak signal gains influence.

Example numbers

For the pair Qwen2.5-7B-Instruct (teacher) to Llama-3.2-3B-Instruct (student), trained on 20,000 maths and code prompts, the paper reports a full-benchmark average of 26.96% for the undistilled student, 32.86% with strict full-vocabulary distillation, 32.64% with strict top-16, 31.59% for the SimCT baseline, and 49.03% for the teacher itself. So distillation closes part of the gap but not most of it, and the simple strict methods beat the baseline.

Caveats

  • This is a preprint, and the authors' code is on an anonymized repository, which suggests the paper is still under review. The results have not been independently reproduced.
  • The student models are small (3B to 7B) and the tests are on maths and code. The pattern might differ for other tasks or larger models.
  • The authors say they do not fully explain why the span supervision is harmful, and they test only reverse KL, not other divergences.

Why it matters if you are learning AI

The general lesson is bigger than this method: supervision quality can matter more than quantity. More training signal is not automatically better if part of it is noisy or contradictory, a point familiar from our article on dirty data. Distillation is also how many compact models that run on a laptop or a phone are made. If you are new to the ideas, see What Is a Transformer?, and for another recent paper on spending computation wisely, ALoDLM. Our Introduction to Natural Language Processing course covers tokenization and language-model training from scratch.

What to watch next

  • Reproduction on other model families, larger students, and tasks beyond maths and code.
  • An explanation of why the mismatch-span signal conflicts with the strict one.
  • Whether the top-k idea is adopted in open-source distillation toolkits.

Sources: arXiv:2610.08448 (submitted 6 October 2026, cs.CL / cs.AI) and its full-text version. This is an independent explainer, not written by the paper's authors.

Related articles