PostHog Open-Sources Jeeves, a 9B Decision Model That Reasons Before It Picks: Beats Jev on Its Own Tests, but Runs Ten Times Slower

Product analytics company PostHog has published Jeeves on GitHub and Hugging Face, an open decision model that mirrors the interface of TypeSafe's Jev: you send a state and yes/no, multiple-choice or rating questions, and get a probability for each option. Unlike Jev, which answers directly, Jeeves writes a reasoning chain for each question before answering. It is built on Qwen3.5-9B with LoRA and a pointer head that scores the options, trained first with supervised fine-tuning on 19,126 questions from 12 public datasets, then with CISPO reinforcement learning, and finally a fitted temperature for probability calibration. The repository reports 0.889 accuracy on its out-of-domain test set versus 0.857 for Jev, and 0.935 on the JevBench public tier versus 0.866, but 0.793 on MMLU knowledge questions versus Jev's 0.900. On a single H100 with full reasoning, median latency is 3.3 seconds with a p90 of 17.1 seconds; with reasoning off it is about 0.3 seconds. The code is MIT licensed, the weights carry Qwen3.5-9B's Apache-2.0 license according to the model card, and the interface is a drop-in replacement for Jev's Python SDK.

Which gap in Jev it is trying to fill

In the two weeks since Jev launched, independent tests have converged on the same picture: fast, cheap and probabilistic, but only about as accurate as mid-priced LLMs, so many teams put a reasoning model behind it as a fallback. Our explainer, What Is Jev?, also recommends using it as a first gate and escalating uncertain cases to an LLM.

Jeeves moves that step back inside the model: let the same model think before it picks. The repository includes a comparison: the same checkpoint scores 0.804 on the test split with reasoning off and 0.840 with it on. Reasoning is sped up with a diffusion draft model adapted to the Qwen3.5 architecture, which the author says gives about a 1.6x speedup.

How to read the numbers

A few things are worth noting:

  • All results are self-reported. The test set, the JevBench items and the evaluation scripts come from PostHog's own repository, and the comparisons with Jev and Kev-9B don't use exactly the same items, as the repository itself states.
  • It wins on judgment, loses on knowledge. Trailing Jev by about 10 points on MMLU shows that a 9B base puts a hard ceiling on knowledge.
  • Latency moves by an order of magnitude. Jev's vendor claims 70 to 500 milliseconds; Jeeves with reasoning has a 3.3-second median and a 17-second tail. It suits offline batches and review passes, not real-time paths where a user is waiting.

What it means for developers

Jeeves's biggest value isn't the score but the fact that you can host it yourself. Jev is only available as a hosted API, so data has to go to TypeSafe's service; Jeeves's weights and training code are public, so it can run on your own GPUs and be trained further on your own data. For teams whose data can't leave their infrastructure, or who need to audit the reasoning, it is currently the only Jev-style option.

The requirements are real too: it needs a CUDA GPU, FP8 acceleration requires a newer architecture, and the weights have to be loaded with Jeeves's own code rather than transformers.

Scope note: all accuracy and latency figures come from the Jeeves repository and model card and have not been independently reproduced; for Kev-9B's origin and settings, see the repository.

via: PostHog/jeeves GitHub repository, Hugging Face model card, AI Weekly report