Deemwar · decision model Live checkGamesDeployTraining bookRecipesDocs
Benchmark · measured 1 Oct 2026

Jeeves needs an H100.
OpenJevX needs a laptop.

We ran OpenJevX (598 MB, CPU only) on the open JevBench public tiers, with calibration, next to the Laya base model it came from. Then we put the published numbers for the big GPU models beside ours. We lose on accuracy, and on the hard tier we are no better than chance. This page says so first.

58.4%for the OpenJevX v0.4.0 model on all 231 public items (95% interval 52 to 65%). The shipped model is now 0.5.2 and has not been re-run on JevBench yet. Easy 100%, original 71%, hard 32.4% against a 33.6% chance rate.
0.25 to 0.34 smedian time per decision, one at a time, on an Apple M5 Pro CPU that was busy with other work. p95 0.71 to 0.80 s.
598 MBmodel file; 1.7 GB RAM at peak. Jeeves 9B needs 21 GB of weights in bf16 and an H100 for its 0.3 s.
ECE 0.121calibration error, against 0.207 for the Laya base. Confidence is partly meaningful, not reliable on hard items. Details below.
1 · Our measurements

Sub-1 GB, CPU: OpenJevX vs Laya base

The open JevBench harness (MIT) by fstandhartinger, commit bb05a335, its own adapter, the 231 public items PostHog's Jeeves report uses: easy 48 + original 72 + hard 111. Three full runs per system, strictly one item at a time, results identical across runs (latency varies). Both models are 8-bit ONNX on CPU.

SystemEasyOriginalHardAll 231ECEBrierp50p95SizePeak RAM
OpenJevX v0.4100%70.8%32.4%58.4%0.1210.5150.25 to 0.34 s0.71 to 0.80 s598 MB1,679 MB
Laya base 421M95.8%70.8%25.2%54.1%0.2070.6080.26 to 0.31 s1.61 to 1.69 s598 MB1,758 MB

Chance rates: easy 28.4%, original 31.1%, hard 33.6%. Lower ECE and Brier are better. Laya base is the unmodified base model that OpenJevX was fine-tuned from, 8-bit quantized by the same script; we did not check it against the upstream weights, and Laya is described by its makers as a base to specialise, not a zero-shot engine.

Where we lose, plainly. OpenJevX beats its own base model by 4 points overall and ties it on the original tier, but on the hard tier its 32.4% (24 to 42%) is not better than the 33.6% you get by guessing. The hard-tier gap to Laya (25.2%) is inside the noise at 111 items. 57 of the 111 hard items exceed OpenJevX's 512-token input cap, but accuracy on the 54 that fit is 31.5%, so truncation does not explain it.

Accuracy by tier, with published GPU numbers on the same items

Accuracy by tierEasy (48)OpenJevX: 100.0%OpenJevX 100.0%Laya base: 95.8%Laya base 95.8%Original (72)OpenJevX: 70.8%OpenJevX 70.8%Laya base: 70.8%Laya base 70.8%Hard (111)OpenJevX: 32.4%OpenJevX 32.4%Laya base: 25.2%Laya base 25.2%Hosted Jev*: 73.0%Hosted Jev* 73.0%Jeeves 9B*: 86.5%Jeeves 9B* 86.5%All 231OpenJevX: 58.4%OpenJevX 58.4%Laya base: 54.1%Laya base 54.1%Hosted Jev*: 86.6%Hosted Jev* 86.6%Jeeves 9B*: 93.5%Jeeves 9B* 93.5%
OpenJevX (measured)Laya base (measured)Hosted Jev *Jeeves 9B *Line = 95% interval · red dashes = chance

* Published by PostHog in the Jeeves model card on the same 231 items (overall 0.866 and 0.935, hard 0.730 and 0.865); we did not run them. Not published for easy or original.

2 · Calibration

Is the confidence number real?

Partly. A confidence of 0.9 or more was right every time, but that band is 22 answers, all easy ones. Below it the number is a rough ranking, and on hard items it carries almost no signal.

Reliability diagramObserved accuracy against stated confidence for OpenJevX and Laya base on 231 JevBench public items. The diagonal is perfect calibration; points below it are overconfident.00252550507575100100stated confidence (%)observed accuracy (%)n=7, confidence 0.27, accuracy 0.00n=13, confidence 0.36, accuracy 0.38n=21, confidence 0.45, accuracy 0.24n=26, confidence 0.55, accuracy 0.46n=26, confidence 0.66, accuracy 0.31n=27, confidence 0.75, accuracy 0.56n=22, confidence 0.84, accuracy 0.45n=89, confidence 0.97, accuracy 0.79n=5, confidence 0.26, accuracy 0.20n=21, confidence 0.35, accuracy 0.19n=18, confidence 0.44, accuracy 0.39n=28, confidence 0.55, accuracy 0.57n=35, confidence 0.65, accuracy 0.43n=45, confidence 0.76, accuracy 0.56n=57, confidence 0.86, accuracy 0.79n=22, confidence 0.93, accuracy 1.00
OpenJevXLaya baseDashed = perfect calibration
OpenJevX confidenceAnswersRight
0.9 and above22100%
0.8 to 0.95779%
0.6 to 0.88050%
below 0.67239%

Laya base for comparison: 111 answers at 0.8 or more, 31 of them wrong; OpenJevX 79 at 0.8 or more, 12 wrong (9 of those 12 on the hard tier).

This is the measurable form of a fair critique: one independent test (Kanta Hayashi, "Jev Does Not Play Dice") found a hosted decision model giving 83% probability to an outcome that came up 19% of the time on a hidden fair die. Our advice is the same as theirs: treat the number as a score, and measure it on your own labelled data before you trust it. That is what our pilot does.

2b · Controls

Is it judging, or is it a prior?

On the easy and original tiers the model beats every control we threw at it, and none of the 231 items appear in OpenJevX's training data (checked, see below). On the hard tier it does not beat the controls: its headroom there is about zero, so its answers are mostly a prior.

Training overlap, checked (2 Oct 2026): we compared every public item against every OpenJevX training file we hold, about 8.7 million rows, by exact match, normalised text and near-duplicate shingles down to Jaccard 0.6. 0 of 231 items matched, and none came within 0.5. The check finds 20 of 20 training rows planted as fake items. Limits:
  • The exact training sample for the v0.4.0 model is no longer available, so we scanned the superset of all local training files instead.
  • The base model's own pre-training data is not public.
  • 60 items share their question wording (not their facts) with a larger internal dataset that OpenJevX's training did not use.
Restricted to the 171 items whose wording appears nowhere, the easy and original headroom stays the same.

A confident answer can be meaningless. An independent study of a hosted decision model found questions that scored well only because of a plain feature in the data (digit count, a short word list) or because the model returned the same answer whatever the input. We ran the same controls on OpenJevX over the same 231 public items (2 Oct 2026). Headroom is |AUC of the real model - 0.5| minus the best control's |AUC - 0.5|; zero or below means the task is not demonstrated, whatever the absolute score. Intervals are bootstrap 95% over items. Aggregate only; no item text is published.

TierItemsModel AUCBest control (|AUC-0.5|)Headroom (first estimate)Headroom, corrected for picking the best control
Easy481.0000.172 (bare ids)+0.328 (+0.189 to +0.366)+0.314 (+0.195 to +0.431)
Original720.8790.167 (bare ids)+0.211 (+0.105 to +0.271)+0.270 (+0.170 to +0.346)
Hard1110.5250.113 (a numeric feature)-0.087 (-0.159 to -0.016)-0.014 (-0.066 to +0.050): no headroom
All 2312310.8060.130 (bare ids)+0.175 (+0.119 to +0.218)+0.185 (+0.141 to +0.262)

AUC mixes yes/no, pick-one and rating items (pick-one is pooled one-vs-rest). Easy and original headroom comes from the pick-one items.

Bare-id nulls for yes/no and rating items are not valid nulls (the stem or the ordering still carries the answer), which only makes our headroom conservative.

The model and every control, side by side (|AUC - 0.5|)

TierModelState = "N/A"State = plausible fillerBare option ids (A/B)Per-item shuffleBest plain featureBest word-list matcher
Easy0.5000.0460.0530.1490.0290.094 (length)0.082
Original0.3790.0320.0480.1670.0500.147 (digit count)0.078
Hard0.0250.0520.0370.0660.0500.113 (sum of numbers)0.058
All 2310.3060.0450.0370.1300.0060.066 (last number)0.029

The plain-feature baselines in full (no model, held-out half, mean of 20 splits)

BaselineEasyOriginalHardAll 231
Character length0.0940.0530.0490.024
Token count0.0550.0540.0580.029
Digit count0.0370.1470.0660.036
First number0.0420.0960.0640.024
Last number0.0360.1290.0960.066
Largest number0.0350.1200.0940.049
Smallest number0.0610.1210.0800.061
Sum of numbers0.0360.1290.1130.045
Count of numbers0.0500.1290.0850.043
Top-20 word list0.0820.0760.0520.029
Top-45 word list0.0820.0780.0580.024

The word lists are fit on the training half only and scored on the held-out half. With 48 easy items they see two or three examples per class, so on easy this control was too starved to count: it could not have reproduced the easy result, and a leave-one-out version is being run. On hard, a single number parsed from the item (their sum) beats the model.

Where we lose

Where the easy 100% and the 22-of-22 band come from

Not from a constant-state prior, length or digit count. With the state deleted or replaced, none of the 22 items in the 0.9-and-above band gives the same answer at similar confidence, and that band's accuracy with the state removed is 0.27 and 0.23. Plain length, digit and number features on easy sit at about chance (best 0.27 accuracy). Memorisation of the items is ruled out as far as we can check: none of the 231 items is in the training files we hold (see the note above, with its limits). One thing we have not ruled out: a word-list explanation. The word-list control on easy has about 2 items per answer, too few to count; leave-one-out and 5-fold versions do not beat chance on any tier, which says nothing either way. Only outside labelled data could settle it.

Sanity checks

Limits: one run per arm; pick-one AUC pooled one-vs-rest; token count is a regex, not the model's tokenizer; the baseline floors are biased upward by half-split noise and by taking the maximum over several baselines, so headroom is conservative; the word-list and numeric baselines see few examples on small tiers; the standard and judge tiers and the official composite were not run. You can run the first controls on your own question in the live check after a run: "Is your question actually judging?".

3 · Published numbers, not measured by us

The GPU reference and other small models

Different benchmarks, different hardware: do not read across rows. Each row says what its number is on. Sources and dates are in the last column; all read on 1 Oct 2026.

ModelAccuracy, on whatLatency, on whatSizeSource
Jeeves 9B (PostHog) GPU0.935 on the same 231 JevBench items; hard 0.865. Weakness: MMLU 0.793 vs 0.900 for Jev.about 0.3 s without thinking, 3.3 s median with thinking, one H100 (fp8); p90 17.1 s with full thinking21 GB weights in bf16 (11.5 GB fp8 on Apple silicon)Jeeves model card and README. Weights Apache-2.0, code MIT.
Hosted Jev 1.13 (TypeSafe) API0.866 on the 231 items; hard 0.730 (as reported by PostHog). 76.0% on Bespoke Labs' 13-dataset set.hosted API; not published by TypeSafenot publishedPostHog model card; Ollama blog; Benchmark Heaven
Hobson 2B (Marc Brooker) GPUNo overall accuracy in the post. "Joint first of 30 at 2B or below" on the JevBench v1.4.2 board (author's claim). Easy set Brier 0.009, accuracy 100%.p50 just over 100 ms, p95 under 300 ms, on an RTX 3090. No CPU or browser numbers.about 2B parametersbrooker.co.za, 28 Sep 2026
Nimble 9B (Bespoke Labs, via Ollama) GPU / Mac75.7% on Bespoke's 13 public datasets (3,880 decisions). Not JevBench.91 ms per decision in one Pac-Man demo, M5 Max9.5 GB (5.6 GB q4)Ollama blog, 29 Sep 2026
Tev1 0.8B (Together AI, via Ollama) sub-1 GB63.5% on Bespoke's 13 datasets (Tev1 4B 73.3%). Not JevBench.not published812 MBOllama library
Laya 421M (Convai) sub-1 GB0.766 on Convai's own 2,000-decision benchmark, only after fine-tuning on its training split; 0.362 zero-shot on the same decisions. Not JevBench.193 to 464 ms on CPU; 32.8 to 39.5 ms on a GPU421M parametersLaya model card, laya.convaiinnovations.com
Julia 1 (Supersonic Labs) sub-1 GB73.15% on its own "typed decisions" set of 2,000, its report33 ms Apple M4, 295 ms Intel i5-1235U, 203 ms Android tablet (median)550 MB, 144M parameterssupersoniclabs.ia.br
Jevstiller distillerAnswers 70.7% of Banking77 traffic locally at 99.45% agreement with Jev, the rest goes to Jevabout 15 ms p50 on CPUsmall local encoder + headpypi.org/project/jevstiller

Not an exhaustive list. Several browser-only models published this week (for example nodd, 18 to 24 MB) quote their own latency but no JevBench accuracy that we could verify, so we leave them out rather than guess.

4 · Speed, size, cost

Per millisecond and per dollar

CPU latency we measured

OpenJevX on an Apple M5 Pro CPU, one decision at a time: p50 0.25 to 0.34 s, p95 0.71 to 0.80 s, while the machine was busy (load average 4 to 10), so read it as pessimistic. Laya base: p50 0.26 to 0.31 s, p95 1.61 to 1.69 s. Hard items are slower (about 0.70 s p50) because they are longer.

In a browser tab (WASM)

We measured OpenJevX in a Chromium tab on the same Mac: about 0.46 s per line one at a time, 0.14 s per line in batches of 16, 1 to 3 s to load from local disk, 2.6 to 2.9 GB of tab memory, answers matching the server to 0.0001 on 32 of 32 checks. WebGPU gave wrong answers on this 8-bit model, so it is off. Public in-browser mode is not live yet: the 598 MB file still needs hosting.

Cost per decision

Hosted Jev is billed per input token: $0.042 per million input tokens, output free (TypeSafe's models page). It publishes no per-decision price; Benchmark Heaven estimates $0.032 per 1,000 decisions, about $32 per million. Jeeves needs an H100-class GPU: no published price. OpenJevX has no per-call bill and the marginal cost on a laptop or server you already own is close to zero, but on a slow rented CPU it can cost more per decision than the hosted API: our 8-vCPU 2 GHz VM manages about 0.7 decisions a second.

Accuracy per dollar, honestly

If your decisions look like the easy and original tiers, a 598 MB model on a CPU you own can answer most of them with no per-call cost and nothing leaving your network. If they look like the hard tier, today's OpenJevX is the wrong tool: use a larger model, or train on your own data and measure it first.

5 · What this does not show

Limits and how to check us

Want it measured on your own decisions?

A public benchmark tells you little about your tickets. A pilot trains on your data and reports, per decision, right and confident, confidently wrong, and abstained, on your own held-out cases.

Email [email protected]

Try the live check · Training book · 18 recipes with real outputs