Open to workNew YorkGet in touch

Decision Models

What the Jeeves decision model's reasoning buys: 5 points for 11x the latency

PostHog's Jeeves adds a reasoning chain to a 9B Jev-compatible classifier. On 325 dev questions, accuracy rises from 0.775 to 0.825. Median latency goes from about 0.3 s to 3.3 s, and p90 reaches 17.1 s.

Published
September 29, 2026
Read
6 min
Author
Samir Sengupta
Chart of Jeeves decision model accuracy versus latency with reasoning off, capped and full on 325 dev questions
Note 072 / 072Daily note · Written from 3 sources

the short version

  • On 325 dev questions, thinking raises the Jeeves decision model from 0.775 to 0.825 accuracy. Median latency rises from about 0.3 s to 3.3 s on one H100, and p90 reaches 17.1 s.
  • The capped setting (max_think 768, nothink_threshold 0.9) scores 0.806 at a 2.0 s median and 5.6 s p90. It keeps 3.1 of the 5.0 accuracy points and cuts p90 from 17.1 s to 5.6 s.
  • With thinking on, Jeeves' largest margins over Jev are on JevBench hard (0.865 vs 0.730), PAWS (0.875 vs 0.788) and held-out rule structures (1.000 vs 0.885). It still trails Jev on knowledge questions (MMLU 0.793 vs 0.900).
  • Jeff's 0.8B model decides in about 22 ms from a single forward pass, but it scores 47.6 on JevBench hard against Jev's 73.3, so the two models serve different latency budgets.

Reasoning does improve the Jeeves decision model. On 325 dev questions, letting it think before it answers raises accuracy from 0.775 to 0.825, and on the 2,962-item test split the same checkpoint goes from 0.804 to 0.840. The cost is latency: median rises from about 0.3 s to 3.3 s on one H100, and p90 reaches 17.1 s. A capped setting recovers most of the gain for less. With max_think 768 and nothink_threshold 0.9, Jeeves scores 0.806 at a 2.0 s median and 5.6 s p90.

PostHog Jeeves is a Qwen3.5-9B model with LoRA and a pointer head, trained with SFT and then CISPO. It accepts Jev's /v1/systemone request format and handles yes/no (noul), multiple-choice (choice) and rating (score) questions in the same request. The repository is MIT-licensed and ships the full training code, train/dev/test data and a block-4 diffusion drafter. The fast comparison point is Jeff, an independent Jev-compatible project whose 0.8B model answers from a single forward pass in about 22 ms on an RTX PRO 6000. Choosing between them comes down to how many seconds per decision a workload can absorb.

How the Jeeves decision model reasons, then decides

The prompt marks the state, question and options with rare, largely unused Qwen tokenizer tokens (<|fim_prefix|>, <|fim_middle|>, <|box_start|>, <|box_end|>, <|fim_suffix|>). The model then writes a reasoning chain. After </think>, the question and options are repeated and a <decide> token is appended. A pointer head scores each option with a scaled dot product. One side is a query projection of the hidden state at <decide>, and the other is a key projection of the hidden state at that option's </opt>. The final probabilities are a softmax over those scores, divided by a temperature fitted on the dev set.

The README reports two ablations that matter to anyone copying the recipe. Replacing the rare tokens with plain text like "State" worsened performance, and so did dropping the repeated question after the reasoning block. SFT ran for 2 epochs (596 steps on 8 GPUs) with LoRA r=16 on all projections, on 19,126 questions from 12 public datasets plus synthetic policy data. Half of those questions carried a reasoning chain sampled from the base model. CISPO followed on 9,992 RL questions, with 8 rollouts each at temperature 1 and a 2,560-token thinking cap. It was stopped at step 402 of a 624-step schedule, because past that point the head over-sharpens on the saturated RL pool.

Where Jeeves beats Jev, and where it trails

All of these figures use thinking with greedy decoding and a 2,560-token cap. On the item-weighted test overall, Jeeves scores 0.889, against 0.822 for Kev-9B and 0.857 for Jev. On JevBench's 231 public items it scores 0.935 against 0.866 for Jev, and on the 111-item public hard tier it scores 0.865 against 0.730. It reaches 1.000 on held-out rule structures and contrastive policies, against Jev's 0.885 and 0.963, and scores 0.875 on PAWS against Jev's 0.788. JevBench calibration error (ECE) on public items is 0.037 for Jeeves and 0.049 for Jev.

The gains are uneven. Jeeves trails Jev on MMLU (0.793 vs 0.900), MMLU-Pro 10-way (0.739 vs 0.840) and transfer overall (0.746 vs 0.800), and slightly on QNLI (0.913 vs 0.925). On unknowable questions answered at p ≥ 0.9, where lower is better, it scores 0.055 against Jev's 0.090 and Kev-9B's 0.000. The README notes that the Kev and Jev comparisons outside JevBench use different items from the same sources. The JevBench rows are restricted to the same public items, and the Kev JevBench figures are for Kev-8B, because no Kev-9B result is published.

Capping max_think cuts p90 latency to 5.6 s

The dev-set table is the clearest statement of the tradeoff. Full thinking gives 0.825 accuracy with 1,138 mean reasoning tokens, at a 3.3 s median and 17.1 s p90. Setting max_think 768 and nothink_threshold 0.9 gives 0.806 with 344 mean reasoning tokens, at a 2.0 s median and 5.6 s p90. No thinking gives 0.775 in about 0.3 s. The capped setting keeps 3.1 of the 5.0 accuracy points and cuts p90 by about two thirds.

  • think (default true): false answers from the prompt alone, in about 0.3 s.
  • max_think (default 2560): truncates each reasoning chain at this many tokens, then answers.
  • nothink_threshold (default null): answers without thinking when the no-think confidence is at least this value.
  • return_reasoning (default false): adds each question's reasoning text to the response.

The README's own example request shows the cost in practice. Three questions with max_think 512, run with FP8 on one H100 and thinking in parallel, produced 1,536 reasoning tokens and a latency of 8,141.6 ms. That is three chains at the 512 cap, and well above the 3.3 s dev median. The diffusion drafter raises chain throughput from 109 tokens per second with plain graphed greedy decoding to 176 with block 4 (1.6x) and 193 with block 8 (1.76x). Eight batched questions at block 4 reach about 960 tokens per second in total. Serving requires Python 3.12 and a CUDA GPU, with Hopper needed for the FP8 kernel. The SDK, a drop-in replacement for Jev's typesafe-sdk, takes max_think and return_reasoning as keyword arguments:

python
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:
    result = client.system_one(
        state="I was charged twice. Please help.",
        questions={
            "billing": Noul(instructions="Is this about billing?"),
        },
        max_think=768,
        return_reasoning=True,
    )

Where Jeff's single forward pass still wins

Jeff returns a calibrated probability per option from one forward pass. It takes about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max via MLX. That is roughly 14 times faster than Jeeves with thinking disabled and about 150 times faster than Jeeves' 3.3 s median with thinking on, although the hardware differs. The price is accuracy on reasoning-heavy items. Jeff-Qwen3.5-0.8B scores 47.6 on JevBench hard and Jeff-Qwen3.5-2B scores 53.3, against Jev's published 73.3 on a 105-item hard tier. Jeeves reports 0.865 on a 111-item hard tier, so the two numbers are not from the same item set.

Jeff's authors say that at this size their models' reasoning won't match Jev's. For cases where zero-shot accuracy is not good enough, they point to a short fine-tune on your own examples. Their voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. For narrow routing tasks with labeled examples, that keeps the 22 ms path. Jeeves' README pitches reasoning at out-of-domain tasks, and its largest reported margins are on JevBench hard and held-out and contrastive policy items.

Hacker News commenters question the 17-second p90

Commenters on Hacker News focused on latency. One asked whether autoregressive reasoning gives up what a Jev-style model buys, meaning a single forward pass and cheap calibrated probabilities. Another asked what the point is at a 17-second p90, saying you might as well use an LLM. A reply pushed back that Jev is not necessarily dirt cheap, reporting that another model was 20% cheaper for their spam detection because of prompt caching. The README keeps the typed probability interface and offers max_think and nothink_threshold to bound the tail, but it still lists 17 s at p90 with full chains as a limitation.

One commenter reported running Jeeves on an M5 Pro with 48GB against a German soccer-tweet irony benchmark. The run took over 30 minutes for 100 tweets and got 68 correct against 79 for Jev, though it beat the other open decision models that commenter had tested. One commenter asked for MPS support, and another asked how Jeeves compares with an ordinary 9B LLM using structured outputs, on both accuracy and speed. A third suggested prompting an LLM to think and then prefilling JSON output instead of post-training. The README does not report a structured-output baseline, so that question stays open.

What is still unknown about Jeeves

The README states that no language consistency reward was used, so the thinking chains are not well interpretable, which limits return_reasoning as an audit trail. The sources report no comparison against a plain 9B LLM with structured outputs and no Jeeves latency figures on hardware other than the H100. The JevBench results exclude the sealed judge tier. The number to watch is how much of the 5.0-point dev gain survives shorter max_think caps on your own traffic. For latency-bound paths, Jeff or Jeeves with think set to false remain the relevant baselines.

Questions this raises

How much slower is the Jeeves decision model with reasoning on?

On one H100, median latency goes from about 0.3 s with thinking disabled to 3.3 s with full thinking, and p90 reaches 17.1 s. Capping max_think at 768 with nothink_threshold 0.9 brings that to a 2.0 s median and 5.6 s p90.

Is Jeeves more accurate than Jev?

It is on several benchmarks, including JevBench public items (0.935 vs 0.866), the public hard tier (0.865 vs 0.730) and PAWS (0.875 vs 0.788). It trails Jev on MMLU (0.793 vs 0.900), MMLU-Pro 10-way (0.739 vs 0.840) and transfer overall (0.746 vs 0.800).

Jeeves vs Jeff decision model speed

Jeff answers from a single forward pass in about 22 ms on an RTX PRO 6000, roughly 150 times faster than Jeeves' 3.3 s median with thinking on, though the hardware differs. Jeff-Qwen3.5-0.8B scores 47.6 on JevBench hard, so it gives up accuracy on reasoning-heavy items.

These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.

More notes

Building something on this?

I ship production LLM, RAG and agentic systems for a living - the infrastructure behind the things these notes are about. Open to roles, contract work and research collaboration.