Skip to content
Every statement on this page is cited and was checked against jevresearch HEAD (4634ffa), openjevx v0.5.9 (ee2a1f4) on 2026-10-04.

Use cases

OpenJevX is built for everyday rules and logs. 1 For rule-shaped decisions and log triage it commits to an answer on 91% of those and is right 92% of the time when it does (model version not recorded). 2 It is not a general reasoner, and larger models do much better on reasoning. 3

Part 1 of the study, inside the company, ran on 2026-10-03; one of the models it compared was a local openjevx on CPU. 4 The source names the local model only as openjevx (CPU, free, private), with no version. 5 The hosted model was Jev jev-1.13.0, paid per call; the text leaves the machine. 6 AUC is how well P(yes) ranks true yes items above true no items: 0.5 is chance and 1.0 is perfect. 7 “Decided” counts items that got a confident yes or no (P ≥ 0.8 or ≤ 0.2). 8 Gold labels were single-pass labels by the researcher (a proxy), or an agent’s own logged classes (“silver” labels). 9

task type sample local openjevx (version not recorded) verdict for local
Is a public comment a criticism of the post? 40 comments AUC 0.76, recall 0.38 local is not good enough to replace the hosted model 10
Can an engineer answer this from experience? 40 AUC 0.63, recall 0.06 Local: not fit 11
A draft-post quality gate: is this a specific first-hand lesson? 40 drafts AUC 0.65, recall 0.06 at 0.8 not fit 12
An email-triage question: does this email need the owner? 50 emails AUC 0.49 (no context), 0.53 (with context) Not fit (local) 13
A lead-quality question 32 AUC 0.59, 2/32 decided Not fit, and labels are too thin 14
Did an agent’s work ship? 53 rows + 26 held-out rows — the hosted produced judgement alone scores AUC 0.837, the same as the averaged score (0.838); adopt produced alone 15
Is an agent idle by design, or idle when it should work? 27 rows local and hosted pick-one: no better nothing works: averaged score 0.41, and a keyword rule’s 67% in-sample equals the constant guess and falls to 53% held out 15

The local model was worse than the hosted model on every task tested: AUC 0.76 vs 0.83, 0.65 vs 0.94, 0.85 vs 0.99. 16 It was also slower in that setup: p50 955 ms per question on CPU against 267 ms hosted. 17 For the hosted model, the draft-post quality gate had AUC 0.94, and all 28 of 40 decided drafts were right. 12

Concrete, high-signal questions work. 18 Fuzzy, judgment-heavy ones fail on the local model: answerable, lead quality, email urgency. 18 The email-triage question scored AUC 0.49 on the local model with no context, which is chance, and 0.53 with context. 13 7 When an answer is unsure, the fix is usually a narrower question: name the exact thing, say what counts as yes, and let code do arithmetic and dates. 19

Good question shapes from the openjevx recipes:

shape decision
Find the failures in a log yes/no per log line: keep only the lines an on-call engineer would act on 20
Route a ticket queue pick-one per ticket: which team owns it 21
Check a claim against the source yes/no: does the document say what you are about to write? 22
Is this command dangerous? yes/no about a command before running it 23
Which tool should handle this? pick-one among function names 24
Several judgements in one call one request, three questions of the three types 25

The recipe outputs come from model 0.5.2 on server 0.5.7, CPU, Apple M5 Pro, run on 2026-10-03, and include the wrong and unsure answers. 26 On that model, the due-date pick was wrong, and confidently: it picked the issue date (0.90). 27 On that model, the prompt-injection recipe (“does fetched text try to instruct an AI assistant?”) missed, with P(yes) 0.34. 28 29 Compute candidate dates in code; the model judges text, it does not do date arithmetic. 30

Lesson 2: pick the bar on held-out data, with controls

Section titled “Lesson 2: pick the bar on held-out data, with controls”

A model can rank well while the threshold is wrong. 31 On the pitchy gate, the hosted model had AUC 0.99 but the bar was off by 0.25. 31 Context-free, the ≤ 0.20 bar still blocks 14 of 53 drafts (26%) that are not pitches. 32 Every gate needs 2-3 known-good and known-bad controls, run whenever the bar is set. 33 A plain rule written after reading the rows scored 67% in-sample, which is only an upper bound. 34 Choosing temperatures on the gate would also fit the test. 35 The OpenJevX training recipe says: hold out calibration before training, then fit one temperature per type. 36 Your inputs will score differently: copy the pattern, not the numbers, and test on your own data. 37

See Benchmarks for gate numbers and latency, and the Model card for honest limits.

  1. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L8 ↩

  2. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L36–37 ↩

  3. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L38–39 ↩

  4. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L3–8 ↩

  5. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L8 ↩

  6. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L6 ↩

  7. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L12 ↩ ↩2

  8. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L13 ↩

  9. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L15–17 ↩

  10. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L21 ↩

  11. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L23 ↩

  12. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L24 ↩ ↩2

  13. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L26 ↩ ↩2

  14. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L27 ↩

  15. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L29 ↩ ↩2

  16. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L39–40 ↩

  17. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L41 ↩

  18. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L50 ↩ ↩2

  19. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L84–85 ↩

  20. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L56 ↩

  21. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L57 ↩

  22. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L60 ↩

  23. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L62 ↩

  24. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L71 ↩

  25. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L72 ↩

  26. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L45–46 ↩

  27. openjevx @ v0.5.9 (ee2a1f4) · recipes/04-pick-a-value.md L48 ↩

  28. openjevx @ v0.5.9 (ee2a1f4) · recipes/08-prompt-injection.md L1–3 ↩

  29. openjevx @ v0.5.9 (ee2a1f4) · recipes/08-prompt-injection.md L49 ↩

  30. openjevx @ v0.5.9 (ee2a1f4) · recipes/04-pick-a-value.md L44 ↩

  31. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L43–44 ↩ ↩2

  32. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L25 ↩

  33. jevresearch @ 4634ffa (4634ffa) · jevresearch/USECASES.md L46 ↩

  34. jevresearch @ 4634ffa (4634ffa) · jevresearch/STATE.md L54 ↩

  35. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L32–33 ↩

  36. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0002-training-run.md L15 ↩

  37. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L50 ↩