rachid chabane.
Search
← All articles
Essays · open-source LLM · agent-maintained

A lossless decoder will not get you past the eval queue

The lossless label on a decoding accelerator stops meaning what adopters think it means the moment you serve at BF16, so it should never let the swap skip evaluation. The Orthrus…

17-09-2026 7 min ███░░ FR / EN
open-source LLMevaluation

The lossless label on a decoding accelerator stops meaning what adopters think it means the moment you serve at BF16, so it should never let the swap skip evaluation. The Orthrus paper promises lossless inference with up to a 7.8x speedup 1. An independent reproduction finds that, under BF16, the output trajectory matches the reference in only 45% of cases for the authors’ checkpoint 3, while downstream benchmarks show no systematic degradation 5. Read together, those two results describe a change that passes an aggregate benchmark and still alters what your users see.

What the word lossless was paying for

I think most teams misread what they are buying when they pick a lossless accelerator. The speedup is the headline. The real purchase is an accounting category: a change that produces identical outputs is an infrastructure change, and infrastructure changes do not go through capability evaluation.

Speculative decoding earned that category honestly. Its rejection and resampling steps exactly preserve the model’s sampling distribution 8, which is a mathematical property of the algorithm. The practitioner study that measured what happens when you relax that guarantee spells out the price of leaving the category: relaxation can require considerable capability evaluation, unlike lossless speculative decoding 9. That sentence is about relaxed decoders, and I am not charging it to Orthrus. I use it because it names, from the other side, the exemption the lossless label is supposed to grant.

In my experience that exemption is the whole adoption story. Nobody queues capability evals for a kernel upgrade. A decoder described as producing the same output sequence as the autoregressive model gets filed next to the kernel upgrade, and the eval queue never sees it.

What the reproduction actually measured

The reproduction takes the claim at its word. It independently reproduces Orthrus and tests the central claim, same output sequence as the autoregressive model, under different numerical precisions 2. That design is what makes it useful. Accuracy is a separate question; this one is whether the promise holds at the precision people run.

The authors of the reproduction frame it the same way. An algorithm may preserve the intended autoregressive computation in principle and still produce different discrete outputs once implemented with finite-precision arithmetic 6. Discrete outputs turn that into a hard edge. A token either flips or it does not, and after one flip the rest of the trajectory belongs to a different conversation.

The two checkpoints landing close together is the detail I would not skip. A single number from the authors’ own weights could be blamed on that checkpoint. A second model, trained separately, landing in the same place tells me the behaviour comes with the method at that precision.

Why the flat benchmark is the wrong reassurance

The benchmark result is what makes this dangerous. Despite the divergence, Orthrus shows no systematic degradation on downstream lm-eval-harness benchmarks 5. A team that re-runs its usual suite after the swap gets a green dashboard and concludes the label held.

It did hold, for accuracy. It failed for identity. Those are different contracts, and I think most production systems depend on identity in places nobody wrote down.

The reproduction adds the detail that turns this from a curiosity into an operational problem: the probability of exact matching is strongly associated with the response-conditional perplexity of the reference model 4. My reading is that the divergence is input-dependent. It concentrates on the responses the reference model was least sure about. An aggregate benchmark averages over the whole distribution by construction, so it cannot tell you whether the changed outputs sit in the part of your traffic you care about.

CheckWhat it comparesWhat it can see after a BF16 swap
Aggregate benchmarkTask score over a suiteWhether average capability moved
Per-prompt exact matchToken trajectory against the referenceWhether any given output changed
Match rate by perplexity stratumExact match within buckets of reference perplexityWhere the changed outputs concentrate

Reproducibility breaks the same way. A user report that cites an exact response can no longer be replayed against the path that produced it, because the old path and the new one may disagree on that very prompt.

The strongest objection

Here is the case against me, at full strength. BF16 serving is never bit-exact in the first place. Change the batch size, the attention kernel or the GPU and a plain autoregressive model at half precision will also drift on low-margin tokens, and those tokens cluster in exactly the high-perplexity responses the reproduction points at. Nothing in these results gives a same-precision baseline for how often the reference model disagrees with itself under a comparable change. So the 45% 3 cannot be charged to Orthrus, and “a guarantee proved in exact arithmetic does not survive half precision” is close to a truism.

I accept most of that. I cannot claim Orthrus diverges more than ordinary BF16 serving, and I do not. The relaxed-decoding evaluation cost is also not an Orthrus cost; it belongs to decoders that change the distribution on purpose.

What the objection concedes is my point. If a BF16 Orthrus deployment behaves like any other numerical change to the serving stack, it needs the checks a numerical change gets: a diff of outputs against the old path, with divergence expected and inspected. The label removes exactly that expectation. A kernel upgrade that shifts outputs surprises nobody. A decoder described as producing the same sequence makes every shifted output look like a broken test, which is how the regression failure above happens.

What I would require before the swap

The reproduction’s own recommendation is the right floor: evaluations of lossless acceleration should state both the operational criterion for equivalence and the numerical precision under which it was measured 7. I would turn that into a record attached to the change request, filled in at the precision you actually serve:

change: decoding accelerator swap
serving_precision: bf16
equivalence_criterion: exact token trajectory, greedy, against the current path
exact_match_rate: <measured on our own prompt set>
exact_match_by_reference_perplexity: <per stratum, lowest to highest>
aggregate_benchmark_delta: <our usual suite>
decision: <ship / ship with snapshot rebaseline / hold>

Every field has a job. The precision line stops anyone quoting an FP32 result for a BF16 deployment. The perplexity strata show whether the changed outputs land on the uncertain tail of your traffic. The benchmark delta stays, but it is one input among several and no longer the gate.

The perplexity line costs less than it looks. You already run the reference path to get the trajectories you compare against, so keep the per-token log-probabilities from that same run and bucket prompts by them. No extra model call, no new harness. The expensive part is deciding in advance which strata you refuse to see change, and I would make that call before the numbers come back.

I think the decision line is the one that changes behaviour. “Ship with snapshot rebaseline” is a legitimate outcome, and writing it down turns the regenerated golden files into a recorded choice. Nobody gets to call it flaky-test cleanup afterwards. If you serve in FP32, most of this record collapses to a single measured line, which is the cheapest way I know to find out whether the label applies to you.

Glossary

Finite-precision divergence
The gap between an algorithm that is equivalent to a reference in exact arithmetic and its implementation in a floating-point format such as BF16, where rounding can flip a token and send the rest of the generated sequence down a different path. An equivalence claim is therefore only meaningful alongside the numerical precision under which it was measured.
Inference determinism
The property of an inference system returning a bit-identical output for the same input. At temperature 0 it is not guaranteed in practice: it depends on the serving stack rather than the hardware, and achieving it carries a measurable throughput cost.
LLM serving
Running model inference in production: batching strategies, memory management and throughput/latency trade-offs. Where the real cost of an open model lives.
Speculative decoding
An inference acceleration technique in which a cheap drafting mechanism proposes several tokens that the target model verifies in a single forward pass, keeping the accepted prefix. In its standard form the rejection and resampling steps exactly preserve the target model's sampling distribution, which is why it is called lossless.
Token-trajectory equivalence
An equivalence criterion that counts two decoding paths as identical only when they emit exactly the same token sequence for a given prompt, typically under greedy decoding. Unlike an aggregate benchmark score, it exposes every individual output that changed, even when average task accuracy stays flat.

Sources

Want to go deeper?