Open to workNew YorkGet in touch

LLM Serving

Jevstiller's disagreement bound: what distilling Jev into a local model guarantees

Jevstiller answers confident Jev classification calls locally in about 15 ms. At a 98% target it bounds disagreement with Jev at 2% of all requests, trading 4 to 8 points of coverage for a guarantee that held on 99 of 100 benchmark splits.

Published
September 29, 2026
Read
7 min
Author
Samir Sengupta
Jevstiller proxy routing requests between a 15 ms local classifier and Jev under a 2% disagreement bound
Note 072 / 073Daily note · Written from 2 sources

the short version

  • The guarantee bounds disagreement with Jev over all requests at 95% confidence per model version. It says nothing about accuracy against ground truth.
  • Choosing the loosest confidence threshold that looks under budget broke a 2% budget on 6 to 12 of 20 splits per task, because it selects on noise.
  • Local coverage depends on how consistent Jev is: about 75 to 87% on intent and news tasks, but only 24 to 29% on tweet tasks where Jev is noisy.
  • The bound holds only while calibration rows are a random sample of live traffic, so the permanent 2% audit is what makes a swap safe after day one.

Jevstiller distills Jev into a small local classifier. With probability at least 95% per model version, it guarantees that at least your target share of all requests, for example 98%, get the label Jev would have returned. Across five public tasks and 100 random splits, the bound exceeded a 2% disagreement budget once, at 2.10%. A naive point-estimate threshold exceeded it on 6 to 12 of 20 splits per task. The cost is four to eight points of local coverage, and the guarantee covers agreement with Jev, not accuracy.

A Jev classification call takes about 300 ms at any load. For an agent loop, a game tick or anything that classifies and then acts, that is the whole budget. Jevstiller sits in front of those calls as a proxy, released under Apache 2.0. It embeds each request with a frozen bge-small encoder (384 dimensions, ONNX Runtime on CPU). On top of that it trains a multinomial logistic regression head with full-batch Adam and early stopping, by cross-entropy against Jev's full probability distribution over the labels.

What the Jevstiller disagreement bound guarantees

The head is a few hundred kilobytes and trains in seconds on a few thousand rows. It is retrained every 2,000 new Jev answers, versioned, and shadow-tested on live traffic before promotion. A request above the confidence threshold and inside the training distribution, judged by a k-nearest-neighbour out-of-distribution scorer, gets a local answer in about 15 ms. Everything else goes to Jev.

Per task, let c be coverage, the share of requests the local model answers. Let e be the share of those where its label differs from Jev's. Forwarded requests agree with Jev by definition, so system agreement is A = 1 − c·e. A target A* sets a budget β = 1 − A*, which is 2 in 100 at 98%, and the router answers as much as it can while keeping c·e under β.

Two things are deliberately left out. First, the bound says nothing about accuracy against the truth: if Jev is wrong, the local model is wrong the same way. Second, it is not a claim about the local model's accuracy on the requests it chose to answer, which the write-up says is what most tools report. It is a statement about disagreement over all traffic, which can be verified against Jev directly.

Why the loosest passing threshold breaks budgets

The usual recipe holds out data, sweeps a confidence threshold, and ships the loosest one whose measured disagreement is within budget. Measured disagreement at any threshold is a noisy estimate. The loosest threshold that looks under budget is therefore preferentially one whose noise pointed downward, and on new traffic it can deliver 2.7% where it measured 1.9%. The benchmark compared that rule against the bound across twenty train/calibration/test splits per task, and the point estimate overshot by up to a full percentage point.

  • Banking77 (77 intents): point estimate 79.8% coverage, mean 1.90%, worst 2.70%, broke budget 9/20; bound 74.9% coverage, mean 1.15%, worst 1.55%, 0/20.
  • CLINC150 (151 intents): point estimate 84.3%, mean 2.09%, worst 2.85%, 12/20; bound 78.9%, mean 1.25%, worst 1.80%, 0/20.
  • AG News (4 classes): point estimate 90.3%, mean 2.06%, worst 2.65%, 11/20; bound 86.5%, mean 1.34%, worst 1.75%, 0/20.
  • TweetEval sentiment: point estimate 28.2%, mean 1.99%, worst 2.60%, 8/20; bound 24.0%, mean 1.46%, worst 2.10%, 1/20.
  • TweetEval offensive: point estimate 36.8%, mean 1.81%, worst 2.70%, 6/20; bound 28.7%, mean 1.10%, worst 1.40%, 0/20.

The point estimate is not a statement at all.
Project write-up, jevstiller.pages.dev

The four rules the bound rests on

  1. The loss is the contract. Each calibration row is IID, labelled by Jev and never used for training. It scores 1 if the local model would answer it and disagree with Jev, else 0. The mean is binomial, so it gets an exact Clopper–Pearson bound with no large-sample approximation.
  2. Thresholds come from a grid fixed before any calibration row is seen. Because the rate can only grow as the threshold loosens, candidates are tested strictest first and testing stops at the first failure. This is fixed-sequence testing as in Learn Then Test. It controls the chance of a wrong choice at 5% with no correction for testing many candidates.
  3. Nothing else touches the calibration rows. The out-of-distribution cutoff is set from the training data, and the candidate thresholds are a fixed grid.
  4. Headroom: the threshold is fitted at 85% of the budget, then rechecked at the full budget on calibration rows pooled with fresh shadowed traffic. The write-up says this recheck is not an independent guarantee, but it catches a threshold sitting exactly on the line.

The bound is a 95% statement, so about 5 in 100 splits may miss, and one miss in 100 is within that. Each retrain spends the 5% again, which is why promotion-time checks are not enough on their own. The write-up does not state a minimum number of calibration rows or the size of the IID calibration split. It says only that the head trains on a few thousand rows and that the status report is worth reading after a few thousand requests. A prior-art review found two errors in the first version: it bounded the selective rate while treating estimated coverage as exact, and it kept the loosest of 200 individually passing thresholds; fixing both moved Banking77 replay coverage only from 70.6% to 70.7%.

When the local model can stand in for Jev

Coverage follows Jev's consistency, not the task's difficulty. On the two tweet tasks, Jev agrees with human labels only 64% and 74% of the time. The local model cannot reproduce that noise within 2%, so under the bound it answered 24.0% and 28.7% locally, against 74.9% to 86.5% on the intent and news tasks. Training on Jev's full probability distributions rather than its top label buys back two to three points on the many-class tasks. Before swapping in, check these points:

  • If your code treats low Jev confidence as unsure, set confidence_floor, available since 0.4.0. Without it, a check at 0.6 lost 8 to 37% of the flags Jev would have raised. With it, those requests go to Jev, at a cost of 4 to 12 points of coverage at a 0.6 floor. It is off unless you set it.
  • Keep the audit on. A fixed 2% of requests always goes to Jev and gives an unbiased interval on live agreement per version. When agreement drops below target the audit rate rises; when the audit's bound confirms a breach, all traffic falls back to Jev and training restarts.
  • Drift recovery was shown in one 24-hour soak, where a stand-in Jev silently changed every answer at hour twelve. The local share fell from 90% to 9% within four minutes. It returned to 90% within 49 minutes, trained on post-change answers only.
  • The target is a dial. Moving from 98% to 95% roughly doubled local answers on the tweet tasks and lifted intent tasks from about 70% to the low 80s. From 99% down to 90%, accuracy against dataset labels stayed within a point of Jev's on these five tasks, which the write-up says is not a law.
  • The contract is per request, not per class, so a rare class can carry most disagreements. Per-class budgets are only on the roadmap.

What Hacker News commenters are pushing back on

One commenter asked whether this is a near drop-in that shrinks a Jev bill over time. The author's opening post notes that it speaks Jev's API only, with an OpenAI-compatible front on the roadmap. Deployment is a Docker container that TYPESAFE_BASE_URL points at. Another commenter argued it only works for simple text classification. A reply countered that it targets questions asked many times, and that the replier's own Jev questions repeat thousands of times.

One commenter asked what type of head this is. The write-up answers that: a multinomial logistic regression layer on frozen bge-small embeddings. The same commenter also wanted a way to correct Jev's wrong answers. Nothing in the write-up describes a correction path, and the contract as written reproduces Jev, errors included. Another commenter contrasted it with model2vec, which distills a sentence transformer into a faster general-purpose encoder, whereas Jevstiller keeps the encoder frozen and distills Jev's decisions on one question into a head.

The benchmarks rerun without an API key, from Jev's recorded answers:

bash
bash experiments/reproduce.sh          # Banking77 result, ~10 min
bash experiments/bench.sh --no-record  # all five tasks and the threshold-rule comparison, an hour or two

What the sources do not report

The evidence covers five public classification tasks and one simulated teacher change in a 24-hour soak. The sources give no minimum calibration size per task and no picture of how the bound tightens as rows accumulate. They report no results on private or multi-turn workloads, and per-class guarantees do not exist yet. The bound assumes calibration rows are a random sample of the traffic they gate, so a shifted mix is left to the audit to catch. The status report prints both the promotion-time bound and the audit's measured interval, and those are the two numbers to watch.

Questions this raises

What does Jevstiller guarantee?

With probability at least 95% per model version, at least your target share of all requests, for example 98%, get the label Jev would have returned. It is a bound on disagreement over all traffic, not on accuracy against the truth. If Jev is wrong, the local model is wrong the same way.

How much coverage does Jevstiller give up compared to a point-estimate threshold?

The bound costs four to eight points of local coverage. On Banking77, for example, coverage fell from 79.8% to 74.9%, while worst-case disagreement fell from 2.70% to 1.55%. The point-estimate threshold broke the 2% budget on 9 of 20 splits, and the bound broke it on none.

Why does Jevstiller answer fewer requests locally on tweet tasks?

Coverage follows Jev's consistency, not task difficulty. Jev agrees with human labels only 64% and 74% of the time on the two TweetEval tasks. As a result, the local model answered only 24.0% and 28.7% of requests, against 74.9% to 86.5% on the intent and news tasks.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.