Open to workNew YorkGet in touch

Decision Models

What the Jeff decision model gives up for 30 ms, and when it can stand in for Jev

Jeff, an open 0.8B Jev-compatible classifier trained on one workstation GPU, decides in 22 to 28 ms but trails Jev's published scores by 16 to 30 points on reasoning-heavy benchmarks.

Published
September 28, 2026
Read
6 min
Author
Samir Sengupta
Jeff 0.8B decision model latency of 22 to 28 ms compared with Jev scores on classification and reasoning benchmarks
Note 070 / 071Daily note · Written from 2 sources

the short version

  • Jeff-Qwen3.5-0.8B beats Jev's published scores on Financial PhraseBank and RAGTruth. It trails by 22 to 30 points on BBH, WinoGrande and JevBench hard.
  • The 0.8B model decides in 22 ms on an RTX PRO 6000 and 28 ms on an M4 Max. Published Doom runs put Jev at 114 to 212 ms per call over the network.
  • Jeff is sensitive to option wording and does no better than random when asked to forecast, so the calling code has to spell out consequences.
  • A short fine-tune moved one voice-navigation task from 31.7% to 95.8% held-out accuracy. That is the README's main path when zero-shot accuracy falls short.

Jeff-Qwen3.5-0.8B is an open-weight decision model that takes Jev's request format and returns calibrated option probabilities. It takes about 22 ms on an RTX PRO 6000 and 28 ms on an Apple M4 Max. In exchange it scores 79.1 against Jev's published 83.0 on a five-benchmark panel, and 47.6 against 73.3 on JevBench's hard tier. It is also sensitive to how options are worded, and does no better than random when asked to forecast. It can replace a hosted Jev call when the decision is classification or grounding over options your code describes by their consequences, and it should not replace one when the decision needs reasoning.

The project is firelex/jeff on GitHub. It fine-tunes Qwen3.5-0.8B, Qwen3.5-2B and Gemma 4 E2B, with training code that starts from the open-source AutoJev recipe. Training ran on a single RTX PRO 6000, where the 0.8B took about 2 hours. The synthetic training data was written by Qwen3.8-Flash-Next on two DGX Sparks. The server exposes a /v1/systemone endpoint that takes the same request format as TypeSafe's Jev, with choice, noul (yes/no) and score question types.

Jeff is not affiliated with or endorsed by TypeSafe. Code is MIT and model weights are Apache 2.0. The training data is not released, and some of its sources are share-alike (CC BY-SA).

Where Jeff matches Jev and where it trails

The README gives a per-benchmark comparison, and the averages hide the split. Jeff-Qwen3.5-0.8B beats Jev's published figures on Financial PhraseBank (96.4 vs 77.0) and RAGTruth (86.1 vs 77.3). It falls well behind on BBH (64.0 vs 94.3), JudgeBench (62.6 vs 78.6), WinoGrande (68.6 vs 90.7) and JevBench hard (47.6 vs 73.3). The authors say Jeff's overall score comes from classification and grounding, and that on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below the large models. The published Jev and AutoJev-27B figures were measured on a different sample of the same benchmarks, so no row is a same-sample comparison.

Fine-tuning accounts for the gain over the base model. The untrained Qwen3.5-0.8B scores 45.3 overall and 36.2 on JevBench hard. The recipe is full-weight fine-tuning for one epoch in batches of 256, with cross-entropy over the option letters, followed by one fitted temperature for calibration. The 2B reaches 83.1 overall, roughly level with Jev's 83.0. Even so, the README recommends the 0.8B for fast option picking, because the 2B is more cautious and plays the game tests worse.

Robustness depends on how the options are worded

Robustness is the less visible cost. The game tests (20 episodes each, seed 1234) show that Jeff works when each option states its consequence in words, such as "you would be hit by a car and lose a life". Asked to forecast instead ("a car arrives in 2 turns"), it does no better than random. Small wording changes move results: giving Frogger's goal option the same words as every other forward option took one episode from 15 crossings to 23. Jev's own Doom prompt, a raw bearing number plus an aiming rule, does not work for any of the Jeff models.


Reason in code, decide with Jeff. It's a classifier, not a planner.
Jeff README, github.com/firelex/jeff

The authors also warn that benchmark scores don't predict game play. The untrained Gemma 4 E2B beats the untrained Qwen models on benchmarks but plays worst. Training fixed its Pac-Man (3.2 to 53.2 pellets) but not its Doom or Frogger. Jeff-Qwen3.5-0.8B matched the hand-coded rule bot on Doom (6.55 kills) and Frogger (10.3 vs 10.25 crossings) but reached only 57.0 of 98 Pac-Man pellets against the bot's 94.1. For reference, Jev's published Doom figure is 6.55 kills when told the aiming rule and -0.60 without it.

What 22 to 28 ms changes in an agent loop

The speed figures are medians over 200 benchmark questions of about 200 input tokens each, run one at a time. Jeff-Qwen3.5-0.8B takes 22 ms on an RTX PRO 6000, 28 ms on an M4 Max with MLX and 463 ms on a 32-thread CPU. Its weights are 1.7 GB at 16-bit, and the 2B's are 4.2 GB, at 24 ms, 60 ms and 708 ms on the same three setups. In the Doom runs, the 0.8B decided in 29 to 49 ms per move on an M4 Max. The README cites Jev's published Doom runs at 114 to 212 ms per call, including the network, and notes the two were not measured on the same hardware.

In a loop that makes many small routing or gating decisions per turn, the per-call gap adds up with every decision. Independent questions can go in one request. The README also recommends short option keys like "1" over long IDs, which cost time and add nothing. On CPU alone, the 463 ms median is slower than Jev's published 114 to 212 ms, though again on different hardware. MLX serving supports the Qwen models only, so Jeff-Gemma4-E2B, at 9.3 GB, has no M4 Max number.

What practitioners on Hacker News push back on

Commenters on Hacker News who tested Jeff against their own tasks reported mixed results. One commenter measured 70% against Jev's 94% on their use cases and called that unacceptable for classification. Another compared Gemini 2.5 Flash Lite, Jev and Jeff on job-ad classification (industry, remote/hybrid/onsite, full or part time). They found the 0.8B not useful, and Jeff-Qwen3.5-2B better but still missing job type.

A reply pointed to fine-tuning as the intended answer. The thread does not show whether the first commenter tried it. The README's own example is a voice-navigation fine-tune on about 11k app-specific examples that took about half an hour on one GPU. It moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision on an M4 Max.

  • One commenter asked why the panel uses reasoning benchmarks like BBH and JudgeBench rather than intent sets such as Banking77 or CLINC150. The sources report no results on either.
  • Another asked how Jeff's speed compares with prompting Qwen to reply with a single token. The README gives no such baseline.
  • One commenter said small task-specific classifiers under 1 MB, trained on CPU in seconds on collected game data, averaged 22 kills on the same Doom scenario. That is a claim from the thread, not a measured comparison in the source.
  • One commenter put Jev's price at $0.42 per million input tokens with output free. The README gives no cost figures for running Jeff.

When a local call can replace Jev

Taken together, the sources support a narrow rule. A hosted Jev call can be swapped for the 0.8B when several conditions hold at once:

  1. The decision is classification or grounding, the families where the README says Jeff matches or beats the large models.
  2. The calling code can describe each option by its consequence rather than asking the model to infer or forecast.
  3. Latency or locality matters, and a GPU or Apple silicon is available, since the 463 ms CPU median removes the speed advantage.
  4. Labelled examples exist to fine-tune if zero-shot accuracy on your own data falls short, as it did for the commenters above.

Keep Jev, or a larger model, for judgement-style calls where BBH, JudgeBench and JevBench hard are the closer proxy. On those three, the 0.8B trails Jev's published scores by 16 to 30 points. Each Jeff answer returns a probability per option, the chosen option and a confidence. That makes it possible to route low-confidence cases to Jev, but the sources do not evaluate such a cascade.

What is still unknown

Several gaps in the published material matter for a production decision. There is no same-sample, same-hardware comparison against Jev, no result on intent datasets like Banking77 or CLINC150, and no single-token-generation speed baseline. There is also no adversarial or context-injection test of Jeff. The authors say the 2B's weaker game play "needs more investigation". Because Jeff and Jev take the same request format, teams considering the swap can run their own labelled set through both before trusting the published averages.

Questions this raises

Can the Jeff decision model replace Jev?

Only for classification or grounding over options your code describes by their consequences. Jeff trails Jev by 16 to 30 points on reasoning-heavy benchmarks such as BBH, JudgeBench and JevBench hard. It also does no better than random when asked to forecast.

How fast is Jeff compared to Jev?

Jeff-Qwen3.5-0.8B takes a median 22 ms on an RTX PRO 6000 and 28 ms on an M4 Max. The README cites Jev's published Doom runs at 114 to 212 ms per call including the network, measured on different hardware. On CPU alone, Jeff's 463 ms median is slower than those Jev figures.

Is Jeff affiliated with TypeSafe's Jev?

No. Jeff is not affiliated with or endorsed by TypeSafe. It exposes a /v1/systemone endpoint that accepts the same request format as Jev, with choice, noul and score question types.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.