We ran OpenJevX (598 MB, CPU only) on the open JevBench public tiers, with calibration, next to the Laya base model it came from. Then we put the published numbers for the big GPU models beside ours. We lose on accuracy, and on the hard tier we are no better than chance. This page says so first.
The open JevBench harness (MIT) by fstandhartinger, commit bb05a335, its own adapter, the 231 public items PostHog's Jeeves report uses: easy 48 + original 72 + hard 111. Three full runs per system, strictly one item at a time, results identical across runs (latency varies). Both models are 8-bit ONNX on CPU.
| System | Easy | Original | Hard | All 231 | ECE | Brier | p50 | p95 | Size | Peak RAM |
|---|---|---|---|---|---|---|---|---|---|---|
| OpenJevX v0.4 | 100% | 70.8% | 32.4% | 58.4% | 0.121 | 0.515 | 0.25 to 0.34 s | 0.71 to 0.80 s | 598 MB | 1,679 MB |
| Laya base 421M | 95.8% | 70.8% | 25.2% | 54.1% | 0.207 | 0.608 | 0.26 to 0.31 s | 1.61 to 1.69 s | 598 MB | 1,758 MB |
Chance rates: easy 28.4%, original 31.1%, hard 33.6%. Lower ECE and Brier are better. Laya base is the unmodified base model that OpenJevX was fine-tuned from, 8-bit quantized by the same script; we did not check it against the upstream weights, and Laya is described by its makers as a base to specialise, not a zero-shot engine.
* Published by PostHog in the Jeeves model card on the same 231 items (overall 0.866 and 0.935, hard 0.730 and 0.865); we did not run them. Not published for easy or original.
Partly. A confidence of 0.9 or more was right every time, but that band is 22 answers, all easy ones. Below it the number is a rough ranking, and on hard items it carries almost no signal.
| OpenJevX confidence | Answers | Right |
|---|---|---|
| 0.9 and above | 22 | 100% |
| 0.8 to 0.9 | 57 | 79% |
| 0.6 to 0.8 | 80 | 50% |
| below 0.6 | 72 | 39% |
Laya base for comparison: 111 answers at 0.8 or more, 31 of them wrong; OpenJevX 79 at 0.8 or more, 12 wrong (9 of those 12 on the hard tier).
This is the measurable form of a fair critique: one independent test (Kanta Hayashi, "Jev Does Not Play Dice") found a hosted decision model giving 83% probability to an outcome that came up 19% of the time on a hidden fair die. Our advice is the same as theirs: treat the number as a score, and measure it on your own labelled data before you trust it. That is what our pilot does.
On the easy and original tiers the model beats every control we threw at it, and none of the 231 items appear in OpenJevX's training data (checked, see below). On the hard tier it does not beat the controls: its headroom there is about zero, so its answers are mostly a prior.
A confident answer can be meaningless. An independent study of a hosted decision model found questions that scored well only because of a plain feature in the data (digit count, a short word list) or because the model returned the same answer whatever the input. We ran the same controls on OpenJevX over the same 231 public items (2 Oct 2026). Headroom is |AUC of the real model - 0.5| minus the best control's |AUC - 0.5|; zero or below means the task is not demonstrated, whatever the absolute score. Intervals are bootstrap 95% over items. Aggregate only; no item text is published.
| Tier | Items | Model AUC | Best control (|AUC-0.5|) | Headroom (first estimate) | Headroom, corrected for picking the best control |
|---|---|---|---|---|---|
| Easy | 48 | 1.000 | 0.172 (bare ids) | +0.328 (+0.189 to +0.366) | +0.314 (+0.195 to +0.431) |
| Original | 72 | 0.879 | 0.167 (bare ids) | +0.211 (+0.105 to +0.271) | +0.270 (+0.170 to +0.346) |
| Hard | 111 | 0.525 | 0.113 (a numeric feature) | -0.087 (-0.159 to -0.016) | -0.014 (-0.066 to +0.050): no headroom |
| All 231 | 231 | 0.806 | 0.130 (bare ids) | +0.175 (+0.119 to +0.218) | +0.185 (+0.141 to +0.262) |
AUC mixes yes/no, pick-one and rating items (pick-one is pooled one-vs-rest). Easy and original headroom comes from the pick-one items.
| Tier | Model | State = "N/A" | State = plausible filler | Bare option ids (A/B) | Per-item shuffle | Best plain feature | Best word-list matcher |
|---|---|---|---|---|---|---|---|
| Easy | 0.500 | 0.046 | 0.053 | 0.149 | 0.029 | 0.094 (length) | 0.082 |
| Original | 0.379 | 0.032 | 0.048 | 0.167 | 0.050 | 0.147 (digit count) | 0.078 |
| Hard | 0.025 | 0.052 | 0.037 | 0.066 | 0.050 | 0.113 (sum of numbers) | 0.058 |
| All 231 | 0.306 | 0.045 | 0.037 | 0.130 | 0.006 | 0.066 (last number) | 0.029 |
| Baseline | Easy | Original | Hard | All 231 |
|---|---|---|---|---|
| Character length | 0.094 | 0.053 | 0.049 | 0.024 |
| Token count | 0.055 | 0.054 | 0.058 | 0.029 |
| Digit count | 0.037 | 0.147 | 0.066 | 0.036 |
| First number | 0.042 | 0.096 | 0.064 | 0.024 |
| Last number | 0.036 | 0.129 | 0.096 | 0.066 |
| Largest number | 0.035 | 0.120 | 0.094 | 0.049 |
| Smallest number | 0.061 | 0.121 | 0.080 | 0.061 |
| Sum of numbers | 0.036 | 0.129 | 0.113 | 0.045 |
| Count of numbers | 0.050 | 0.129 | 0.085 | 0.043 |
| Top-20 word list | 0.082 | 0.076 | 0.052 | 0.029 |
| Top-45 word list | 0.082 | 0.078 | 0.058 | 0.024 |
The word lists are fit on the training half only and scored on the held-out half. With 48 easy items they see two or three examples per class, so on easy this control was too starved to count: it could not have reproduced the easy result, and a leave-one-out version is being run. On hard, a single number parsed from the item (their sum) beats the model.
Not from a constant-state prior, length or digit count. With the state deleted or replaced, none of the 22 items in the 0.9-and-above band gives the same answer at similar confidence, and that band's accuracy with the state removed is 0.27 and 0.23. Plain length, digit and number features on easy sit at about chance (best 0.27 accuracy). Memorisation of the items is ruled out as far as we can check: none of the 231 items is in the training files we hold (see the note above, with its limits). One thing we have not ruled out: a word-list explanation. The word-list control on easy has about 2 items per answer, too few to count; leave-one-out and 5-fold versions do not beat chance on any tier, which says nothing either way. Only outside labelled data could settle it.
Limits: one run per arm; pick-one AUC pooled one-vs-rest; token count is a regex, not the model's tokenizer; the baseline floors are biased upward by half-split noise and by taking the maximum over several baselines, so headroom is conservative; the word-list and numeric baselines see few examples on small tiers; the standard and judge tiers and the official composite were not run. You can run the first controls on your own question in the live check after a run: "Is your question actually judging?".
Different benchmarks, different hardware: do not read across rows. Each row says what its number is on. Sources and dates are in the last column; all read on 1 Oct 2026.
| Model | Accuracy, on what | Latency, on what | Size | Source |
|---|---|---|---|---|
| Jeeves 9B (PostHog) GPU | 0.935 on the same 231 JevBench items; hard 0.865. Weakness: MMLU 0.793 vs 0.900 for Jev. | about 0.3 s without thinking, 3.3 s median with thinking, one H100 (fp8); p90 17.1 s with full thinking | 21 GB weights in bf16 (11.5 GB fp8 on Apple silicon) | Jeeves model card and README. Weights Apache-2.0, code MIT. |
| Hosted Jev 1.13 (TypeSafe) API | 0.866 on the 231 items; hard 0.730 (as reported by PostHog). 76.0% on Bespoke Labs' 13-dataset set. | hosted API; not published by TypeSafe | not published | PostHog model card; Ollama blog; Benchmark Heaven |
| Hobson 2B (Marc Brooker) GPU | No overall accuracy in the post. "Joint first of 30 at 2B or below" on the JevBench v1.4.2 board (author's claim). Easy set Brier 0.009, accuracy 100%. | p50 just over 100 ms, p95 under 300 ms, on an RTX 3090. No CPU or browser numbers. | about 2B parameters | brooker.co.za, 28 Sep 2026 |
| Nimble 9B (Bespoke Labs, via Ollama) GPU / Mac | 75.7% on Bespoke's 13 public datasets (3,880 decisions). Not JevBench. | 91 ms per decision in one Pac-Man demo, M5 Max | 9.5 GB (5.6 GB q4) | Ollama blog, 29 Sep 2026 |
| Tev1 0.8B (Together AI, via Ollama) sub-1 GB | 63.5% on Bespoke's 13 datasets (Tev1 4B 73.3%). Not JevBench. | not published | 812 MB | Ollama library |
| Laya 421M (Convai) sub-1 GB | 0.766 on Convai's own 2,000-decision benchmark, only after fine-tuning on its training split; 0.362 zero-shot on the same decisions. Not JevBench. | 193 to 464 ms on CPU; 32.8 to 39.5 ms on a GPU | 421M parameters | Laya model card, laya.convaiinnovations.com |
| Julia 1 (Supersonic Labs) sub-1 GB | 73.15% on its own "typed decisions" set of 2,000, its report | 33 ms Apple M4, 295 ms Intel i5-1235U, 203 ms Android tablet (median) | 550 MB, 144M parameters | supersoniclabs.ia.br |
| Jevstiller distiller | Answers 70.7% of Banking77 traffic locally at 99.45% agreement with Jev, the rest goes to Jev | about 15 ms p50 on CPU | small local encoder + head | pypi.org/project/jevstiller |
Not an exhaustive list. Several browser-only models published this week (for example nodd, 18 to 24 MB) quote their own latency but no JevBench accuracy that we could verify, so we leave them out rather than guess.
OpenJevX on an Apple M5 Pro CPU, one decision at a time: p50 0.25 to 0.34 s, p95 0.71 to 0.80 s, while the machine was busy (load average 4 to 10), so read it as pessimistic. Laya base: p50 0.26 to 0.31 s, p95 1.61 to 1.69 s. Hard items are slower (about 0.70 s p50) because they are longer.
We measured OpenJevX in a Chromium tab on the same Mac: about 0.46 s per line one at a time, 0.14 s per line in batches of 16, 1 to 3 s to load from local disk, 2.6 to 2.9 GB of tab memory, answers matching the server to 0.0001 on 32 of 32 checks. WebGPU gave wrong answers on this 8-bit model, so it is off. Public in-browser mode is not live yet: the 598 MB file still needs hosting.
Hosted Jev is billed per input token: $0.042 per million input tokens, output free (TypeSafe's models page). It publishes no per-decision price; Benchmark Heaven estimates $0.032 per 1,000 decisions, about $32 per million. Jeeves needs an H100-class GPU: no published price. OpenJevX has no per-call bill and the marginal cost on a laptop or server you already own is close to zero, but on a slow rented CPU it can cost more per decision than the hosted API: our 8-vCPU 2 GHz VM manages about 0.7 decisions a second.
If your decisions look like the easy and original tiers, a 598 MB model on a CPU you own can answer most of them with no per-call cost and nothing leaving your network. If they look like the hard tier, today's OpenJevX is the wrong tool: use a larger model, or train on your own data and measure it first.
/v1/systemone endpoint with python -m jevbench.cli run --tasks datasets/public/<tier>.jsonl --adapter typesafe --endpoint http://127.0.0.1:<port>. Our aggregate results are in results.json (no item text). Questions about our method: [email protected].A public benchmark tells you little about your tickets. A pilot trains on your data and reports, per decision, right and confident, confidently wrong, and abstained, on your own held-out cases.
Email [email protected]Try the live check · Training book · 18 recipes with real outputs