Skip to content
Every statement on this page is cited and was checked against jev-cloud origin/main (10a744c), openjevx v0.5.9 (ee2a1f4) on 2026-10-04.

Benchmarks

Every number below names its model version and, for latency, its hardware.

The scorer uses jevx defaults: yes at or above 0.8, no at or below 0.2, choice and score at or above 0.6. 1 A release must reach right & confident of at least 90% and confidently wrong of at most 2% on the basics gates. 2 It must also get at least 12 of the 13 jevx fundamentals. 3

gate set (size) accuracy right & confident confidently wrong
Everyday basics (463) 99.1% 98.7% 0.9% 4
Rule-checking basics (300), v0.5.2 100% 100% 0.0% 5
conditions_eval (11,949) 98.9% 98.8% 1.0% 6
it_worker_eval (9,159) 98.5% 97.8% 1.0% 7 8
logs gate (900) 92.0% 91.9% 7.8% 7 9

Model 0.5.2 got 13 of the 13 jevx fundamentals. 7 10 The named miss: on conditions_eval, confidently wrong rose from 101 to 118 of 11,949 (0.85% to 0.99%). 11

On the logs gate (900), model 0.5.2 is 7.8% confidently wrong. 7 9 Known miss: on real incident postmortems it is confidently wrong 18% of the time. 12

Model version not recorded for this figure.

On real public postmortems OpenJevX was confident and right 50% of the time. 13 14 The rules set shares facts with the training data (the questions are reworded), so its 100% overstates performance on new rules. 15 Each of those comparison sets had 100 questions, so each number is within about ±9 points. 16 There, OpenJevX’s median time per question was 450 ms on a CPU shared with another job. 17 18 On a general panel (BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande) it scored 60.4% accuracy and was confident on 11%. 19

Both model folders ran on the same Go server build (ORT 1.29, CPU) with finetuning/gate/gate_eval.py. 20

Hardware not recorded beyond “CPU”.

The postmortems, bigvul, HDFS and BGL sets are scored in full; the others are a fixed 300-row sample, so each of those numbers is within about ±5 points. 21

set (n) v0.5.0 acc / right&conf / conf-wrong v0.5.2 acc / right&conf / conf-wrong
Dan Luu public postmortems (102) 60.8% / 52.9% / 23.5% 68.6% / 62.7% / 20.6% 22
bigvul, CVE fixes (201) 47.8% / 0.0% / 0.0% 45.8% / 0.0% / 0.0% 23 24
defect detection (300) 64.0% / 21.0% / 3.0% 64.0% / 20.3% / 4.3% 23 25
code search (300) 92.7% / 90.0% / 4.0% 94.0% / 90.0% / 4.0% 23 26
ms_marco relevance (300) 69.7% / 31.3% / 8.0% 72.3% / 32.0% / 6.7% 23 27
HDFS operator logs (118) 67.8% / 18.6% / 0.0% 68.6% / 18.6% / 0.0% 23 28
BGL operator logs (214) 76.6% / 73.4% / 21.0% 64.5% / 64.5% / 35.0% 23 29
logs_eval, in-distribution (1,224) 95.4% / 95.3% / 4.5% 95.3% / 95.3% / 4.5% 23 30

Warning: BGL regression in model 0.5.2. On BGL (Blue Gene/L supercomputer log) alerts, v0.5.2 misses alerts it used to catch. 31 Alerts caught: v0.5.0 58, v0.5.2 run 1 30, v0.5.2 run 2 32. 32 Confidently wrong goes from 21.0% to 35.0%, about 3 standard errors. 33 The release gate does not cover BGL: its logs gate is application/database/frontend logs, and it held at 7.8%. 34 The report’s most likely cause is interference from the rule-reading drills. 35

Model 0.5.0 on an Apple M5 Pro, 18 cores, CPU only; times in ms, end to end in process. 36 37 The table shows the “after” columns of Fix 1, ONNX Runtime 1.22 -> 1.29. 38 39

case tokens p50 ms p95 ms
short x1 36 22.4 23.0 40
typical x1 164 84.8 85.8 39 41
long x1, after the fix 912 539.6 576.6 42
typical x8 questions 1319 666.4 689.5 43

Cold start was about 1.0 s once per process (model 0.5.0, M5 Pro). 44 45

Latency in a Linux arm64 container, model 0.5.0

Section titled “Latency in a Linux arm64 container, model 0.5.0”

Same model 0.5.0, Linux arm64 container with --cpus=4, ms p50 / p95. 44 46

case tokens p50 / p95 ms
short x1, after (4 threads) 36 42.2 / 43.3 47
typical x1, after (4 threads) 164 167.4 / 170.9 48
long x1, after (4 threads) 912 1055.5 / 1074.1 49

Model 0.5.2, ONNX Runtime 1.29.0 CPU; short is 40 tokens and long is 512 tokens; median of 3 runs after a warm-up. 50

CPU (threads) short / long, shipped default
AMD EPYC 7763, AVX2 (1) 531 / 6953 ms 51
AMD EPYC 7763, AVX2 (2) 289 / 3600 ms 52
AMD EPYC 9V74, AVX2 (1) 566 / 7445 ms 53
Intel 8573C, AVX-512 VNNI (2) 173 / 2104 ms 54
Apple M5 Pro (1) 109 / 1445 ms 55

On model 0.5.2, x86 AVX2 is about 4.8x slower per core than the M5 Pro. 56 57

Fargate times are server-side run times from the Server-Timing header, model 0.5.2, p50 / p95 ms. 58

task short, 51 tokens long, 512 tokens
1 vCPU, 2 threads (from GOMAXPROCS) 1510 / 1585 16123 / 16277 59
1 vCPU, OPENJEVX_THREADS=1 362 / 374 4263 / 4406 60
2 vCPU, 2 threads 354 / 373 3895 / 4028 61

A 1 vCPU Fargate task sees 2 CPUs, so 2 threads there are throttled and 4.2x slower. 62 Since server 0.5.5 the thread count comes from the cgroup quota or the ECS task’s Limits.CPU. 63

Model 0.5.2 trained on one RTX 4090 on Vast.ai: 791,239 decisions after packaging, one full pass in 2.5 h. 64 65

See the Model card for what the model is, and Use cases for how it does on real tasks.

  1. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L4–5 ↩

  2. openjevx @ v0.5.9 (ee2a1f4) · docs/RELEASE_PROCESS.md L72–78 ↩

  3. openjevx @ v0.5.9 (ee2a1f4) · docs/RELEASE_PROCESS.md L77–78 ↩

  4. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L107–109 ↩

  5. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L107–110 ↩

  6. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L107–111 ↩

  7. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L107 ↩ ↩2 ↩3 ↩4

  8. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L112 ↩

  9. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L113 ↩ ↩2

  10. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L114 ↩

  11. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L116–117 ↩

  12. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L40 ↩

  13. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L10 ↩

  14. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L15 ↩

  15. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L41–42 ↩

  16. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L3–4 ↩

  17. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L28 ↩

  18. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L32 ↩

  19. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L18–23 ↩

  20. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L3 ↩

  21. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L4–6 ↩

  22. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L8–10 ↩

  23. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L8 ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7

  24. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L11 ↩

  25. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L12 ↩

  26. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L13 ↩

  27. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L14 ↩

  28. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L15 ↩

  29. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L16 ↩

  30. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L17 ↩

  31. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L21–22 ↩

  32. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L24–28 ↩

  33. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L30–31 ↩

  34. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L40 ↩

  35. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L36 ↩

  36. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L3–4 ↩

  37. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L39 ↩

  38. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L25 ↩

  39. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L31 ↩ ↩2

  40. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L31–33 ↩

  41. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L34 ↩

  42. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L31–35 ↩

  43. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L31–36 ↩

  44. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L3 ↩ ↩2

  45. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L40 ↩

  46. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L84–85 ↩

  47. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L85–89 ↩

  48. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L85–90 ↩

  49. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L85–91 ↩

  50. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L3–5 ↩

  51. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L26–28 ↩

  52. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L26–29 ↩

  53. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L26–30 ↩

  54. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L26–31 ↩

  55. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L26–32 ↩

  56. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L3 ↩

  57. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L23–24 ↩

  58. jev-cloud @ origin/main (10a744c) · docs/fargate-retest-2026-10-03.md L14–15 ↩

  59. jev-cloud @ origin/main (10a744c) · docs/fargate-retest-2026-10-03.md L17–19 ↩

  60. jev-cloud @ origin/main (10a744c) · docs/fargate-retest-2026-10-03.md L17–20 ↩

  61. jev-cloud @ origin/main (10a744c) · docs/fargate-retest-2026-10-03.md L17–21 ↩

  62. jev-cloud @ origin/main (10a744c) · docs/fargate-retest-2026-10-03.md L23–24 ↩

  63. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L49–50 ↩

  64. openjevx @ v0.5.9 (ee2a1f4) · README.md L224 ↩

  65. openjevx @ v0.5.9 (ee2a1f4) · README.md L239 ↩