Which gap in Jev it is trying to fill
In the two weeks since Jev launched, independent tests have converged on the same picture: fast, cheap and probabilistic, but only about as accurate as mid-priced LLMs, so many teams put a reasoning model behind it as a fallback. Our explainer, What Is Jev?, also recommends using it as a first gate and escalating uncertain cases to an LLM.
Jeeves moves that step back inside the model: let the same model think before it picks. The repository includes a comparison: the same checkpoint scores 0.804 on the test split with reasoning off and 0.840 with it on. Reasoning is sped up with a diffusion draft model adapted to the Qwen3.5 architecture, which the author says gives about a 1.6x speedup.
How to read the numbers
A few things are worth noting:
- All results are self-reported. The test set, the JevBench items and the evaluation scripts come from PostHog's own repository, and the comparisons with Jev and Kev-9B don't use exactly the same items, as the repository itself states.
- It wins on judgment, loses on knowledge. Trailing Jev by about 10 points on MMLU shows that a 9B base puts a hard ceiling on knowledge.
- Latency moves by an order of magnitude. Jev's vendor claims 70 to 500 milliseconds; Jeeves with reasoning has a 3.3-second median and a 17-second tail. It suits offline batches and review passes, not real-time paths where a user is waiting.
What it means for developers
Jeeves's biggest value isn't the score but the fact that you can host it yourself. Jev is only available as a hosted API, so data has to go to TypeSafe's service; Jeeves's weights and training code are public, so it can run on your own GPUs and be trained further on your own data. For teams whose data can't leave their infrastructure, or who need to audit the reasoning, it is currently the only Jev-style option.
The requirements are real too: it needs a CUDA GPU, FP8 acceleration requires a newer architecture, and the weights have to be loaded with Jeeves's own code rather than transformers.
Scope note: all accuracy and latency figures come from the Jeeves repository and model card and have not been independently reproduced; for Kev-9B's origin and settings, see the repository.
via: PostHog/jeeves GitHub repository, Hugging Face model card, AI Weekly report