Benchmarks
Every number below names its model version and, for latency, its hardware.
How the gate scores
Section titled “How the gate scores”The scorer uses jevx defaults: yes at or above 0.8, no at or below 0.2, choice and score at or above 0.6. 1 A release must reach right & confident of at least 90% and confidently wrong of at most 2% on the basics gates. 2 It must also get at least 12 of the 13 jevx fundamentals. 3
Release gate, model 0.5.2
Section titled “Release gate, model 0.5.2”| gate set (size) | accuracy | right & confident | confidently wrong |
|---|---|---|---|
| Everyday basics (463) | 99.1% | 98.7% | 0.9% 4 |
| Rule-checking basics (300), v0.5.2 | 100% | 100% | 0.0% 5 |
| conditions_eval (11,949) | 98.9% | 98.8% | 1.0% 6 |
| it_worker_eval (9,159) | 98.5% | 97.8% | 1.0% 7 8 |
| logs gate (900) | 92.0% | 91.9% | 7.8% 7 9 |
Model 0.5.2 got 13 of the 13 jevx fundamentals. 7 10 The named miss: on conditions_eval, confidently wrong rose from 101 to 118 of 11,949 (0.85% to 0.99%). 11
The misses
Section titled “The misses”On the logs gate (900), model 0.5.2 is 7.8% confidently wrong. 7 9 Known miss: on real incident postmortems it is confidently wrong 18% of the time. 12
Model version not recorded for this figure.
On real public postmortems OpenJevX was confident and right 50% of the time. 13 14 The rules set shares facts with the training data (the questions are reworded), so its 100% overstates performance on new rules. 15 Each of those comparison sets had 100 questions, so each number is within about ±9 points. 16 There, OpenJevX’s median time per question was 450 ms on a CPU shared with another job. 17 18 On a general panel (BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande) it scored 60.4% accuracy and was confident on 11%. 19
Benchmark sets, model 0.5.2 vs 0.5.0
Section titled “Benchmark sets, model 0.5.2 vs 0.5.0”Both model folders ran on the same Go server build (ORT 1.29, CPU) with finetuning/gate/gate_eval.py. 20
Hardware not recorded beyond “CPU”.
The postmortems, bigvul, HDFS and BGL sets are scored in full; the others are a fixed 300-row sample, so each of those numbers is within about ±5 points. 21
| set (n) | v0.5.0 acc / right&conf / conf-wrong | v0.5.2 acc / right&conf / conf-wrong |
|---|---|---|
| Dan Luu public postmortems (102) | 60.8% / 52.9% / 23.5% | 68.6% / 62.7% / 20.6% 22 |
| bigvul, CVE fixes (201) | 47.8% / 0.0% / 0.0% | 45.8% / 0.0% / 0.0% 23 24 |
| defect detection (300) | 64.0% / 21.0% / 3.0% | 64.0% / 20.3% / 4.3% 23 25 |
| code search (300) | 92.7% / 90.0% / 4.0% | 94.0% / 90.0% / 4.0% 23 26 |
| ms_marco relevance (300) | 69.7% / 31.3% / 8.0% | 72.3% / 32.0% / 6.7% 23 27 |
| HDFS operator logs (118) | 67.8% / 18.6% / 0.0% | 68.6% / 18.6% / 0.0% 23 28 |
| BGL operator logs (214) | 76.6% / 73.4% / 21.0% | 64.5% / 64.5% / 35.0% 23 29 |
| logs_eval, in-distribution (1,224) | 95.4% / 95.3% / 4.5% | 95.3% / 95.3% / 4.5% 23 30 |
Warning: BGL regression in model 0.5.2. On BGL (Blue Gene/L supercomputer log) alerts, v0.5.2 misses alerts it used to catch. 31 Alerts caught: v0.5.0 58, v0.5.2 run 1 30, v0.5.2 run 2 32. 32 Confidently wrong goes from 21.0% to 35.0%, about 3 standard errors. 33 The release gate does not cover BGL: its logs gate is application/database/frontend logs, and it held at 7.8%. 34 The report’s most likely cause is interference from the rule-reading drills. 35
Latency on Apple M5 Pro, model 0.5.0
Section titled “Latency on Apple M5 Pro, model 0.5.0”Model 0.5.0 on an Apple M5 Pro, 18 cores, CPU only; times in ms, end to end in process. 36 37 The table shows the “after” columns of Fix 1, ONNX Runtime 1.22 -> 1.29. 38 39
| case | tokens | p50 ms | p95 ms |
|---|---|---|---|
| short x1 | 36 | 22.4 | 23.0 40 |
| typical x1 | 164 | 84.8 | 85.8 39 41 |
| long x1, after the fix | 912 | 539.6 | 576.6 42 |
| typical x8 questions | 1319 | 666.4 | 689.5 43 |
Cold start was about 1.0 s once per process (model 0.5.0, M5 Pro). 44 45
Latency in a Linux arm64 container, model 0.5.0
Section titled “Latency in a Linux arm64 container, model 0.5.0”Same model 0.5.0, Linux arm64 container with --cpus=4, ms p50 / p95. 44 46
| case | tokens | p50 / p95 ms |
|---|---|---|
| short x1, after (4 threads) | 36 | 42.2 / 43.3 47 |
| typical x1, after (4 threads) | 164 | 167.4 / 170.9 48 |
| long x1, after (4 threads) | 912 | 1055.5 / 1074.1 49 |
Latency on x86 and M5 Pro, model 0.5.2
Section titled “Latency on x86 and M5 Pro, model 0.5.2”Model 0.5.2, ONNX Runtime 1.29.0 CPU; short is 40 tokens and long is 512 tokens; median of 3 runs after a warm-up. 50
| CPU (threads) | short / long, shipped default |
|---|---|
| AMD EPYC 7763, AVX2 (1) | 531 / 6953 ms 51 |
| AMD EPYC 7763, AVX2 (2) | 289 / 3600 ms 52 |
| AMD EPYC 9V74, AVX2 (1) | 566 / 7445 ms 53 |
| Intel 8573C, AVX-512 VNNI (2) | 173 / 2104 ms 54 |
| Apple M5 Pro (1) | 109 / 1445 ms 55 |
On model 0.5.2, x86 AVX2 is about 4.8x slower per core than the M5 Pro. 56 57
Latency on AWS Fargate, model 0.5.2
Section titled “Latency on AWS Fargate, model 0.5.2”Fargate times are server-side run times from the Server-Timing header, model 0.5.2, p50 / p95 ms. 58
| task | short, 51 tokens | long, 512 tokens |
|---|---|---|
| 1 vCPU, 2 threads (from GOMAXPROCS) | 1510 / 1585 | 16123 / 16277 59 |
1 vCPU, OPENJEVX_THREADS=1 |
362 / 374 | 4263 / 4406 60 |
| 2 vCPU, 2 threads | 354 / 373 | 3895 / 4028 61 |
A 1 vCPU Fargate task sees 2 CPUs, so 2 threads there are throttled and 4.2x slower. 62
Since server 0.5.5 the thread count comes from the cgroup quota or the ECS task’s Limits.CPU. 63
Training compute, model 0.5.2
Section titled “Training compute, model 0.5.2”Model 0.5.2 trained on one RTX 4090 on Vast.ai: 791,239 decisions after packaging, one full pass in 2.5 h. 64 65
See the Model card for what the model is, and Use cases for how it does on real tasks.
Sources
Section titled “Sources”Footnotes
Section titled “Footnotes”-
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL4–5 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/RELEASE_PROCESS.mdL72–78 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/RELEASE_PROCESS.mdL77–78 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL107–109 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL107–110 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL107–111 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL107 ↩ ↩2 ↩3 ↩4 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL112 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL113 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL114 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL116–117 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL10 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL41–42 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL3–4 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL28 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL32 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL18–23 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL3 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL4–6 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL8–10 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL8 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL11 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL12 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL13 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL14 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL16 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL17 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL21–22 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL24–28 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL30–31 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL36 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL3–4 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL39 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL25 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL31 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL31–33 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL34 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL31–35 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL31–36 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL3 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL84–85 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL85–89 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL85–90 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL85–91 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL3–5 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL26–28 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL26–29 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL26–30 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL26–31 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL26–32 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL3 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL23–24 ↩ -
jev-cloud @ origin/main (10a744c) ·
docs/fargate-retest-2026-10-03.mdL14–15 ↩ -
jev-cloud @ origin/main (10a744c) ·
docs/fargate-retest-2026-10-03.mdL17–19 ↩ -
jev-cloud @ origin/main (10a744c) ·
docs/fargate-retest-2026-10-03.mdL17–20 ↩ -
jev-cloud @ origin/main (10a744c) ·
docs/fargate-retest-2026-10-03.mdL17–21 ↩ -
jev-cloud @ origin/main (10a744c) ·
docs/fargate-retest-2026-10-03.mdL23–24 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL49–50 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL224 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL239 ↩