Use cases
Where it fits
Section titled “Where it fits”OpenJevX is built for everyday rules and logs. 1 For rule-shaped decisions and log triage it commits to an answer on 91% of those and is right 92% of the time when it does (model version not recorded). 2 It is not a general reasoner, and larger models do much better on reasoning. 3
Measured on real task types
Section titled “Measured on real task types”Part 1 of the study, inside the company, ran on 2026-10-03; one of the models it compared was a local openjevx on CPU. 4
The source names the local model only as openjevx (CPU, free, private), with no version. 5
The hosted model was Jev jev-1.13.0, paid per call; the text leaves the machine. 6
AUC is how well P(yes) ranks true yes items above true no items: 0.5 is chance and 1.0 is perfect. 7
“Decided” counts items that got a confident yes or no (P ≥ 0.8 or ≤ 0.2). 8
Gold labels were single-pass labels by the researcher (a proxy), or an agent’s own logged classes (“silver” labels). 9
| task type | sample | local openjevx (version not recorded) | verdict for local |
|---|---|---|---|
| Is a public comment a criticism of the post? | 40 comments | AUC 0.76, recall 0.38 | local is not good enough to replace the hosted model 10 |
| Can an engineer answer this from experience? | 40 | AUC 0.63, recall 0.06 | Local: not fit 11 |
| A draft-post quality gate: is this a specific first-hand lesson? | 40 drafts | AUC 0.65, recall 0.06 at 0.8 | not fit 12 |
| An email-triage question: does this email need the owner? | 50 emails | AUC 0.49 (no context), 0.53 (with context) | Not fit (local) 13 |
| A lead-quality question | 32 | AUC 0.59, 2/32 decided | Not fit, and labels are too thin 14 |
| Did an agent’s work ship? | 53 rows + 26 held-out rows | — | the hosted produced judgement alone scores AUC 0.837, the same as the averaged score (0.838); adopt produced alone 15 |
| Is an agent idle by design, or idle when it should work? | 27 rows | local and hosted pick-one: no better | nothing works: averaged score 0.41, and a keyword rule’s 67% in-sample equals the constant guess and falls to 53% held out 15 |
The local model was worse than the hosted model on every task tested: AUC 0.76 vs 0.83, 0.65 vs 0.94, 0.85 vs 0.99. 16 It was also slower in that setup: p50 955 ms per question on CPU against 267 ms hosted. 17 For the hosted model, the draft-post quality gate had AUC 0.94, and all 28 of 40 decided drafts were right. 12
Lesson 1: crisp, checkable questions work
Section titled “Lesson 1: crisp, checkable questions work”Concrete, high-signal questions work. 18 Fuzzy, judgment-heavy ones fail on the local model: answerable, lead quality, email urgency. 18 The email-triage question scored AUC 0.49 on the local model with no context, which is chance, and 0.53 with context. 13 7 When an answer is unsure, the fix is usually a narrower question: name the exact thing, say what counts as yes, and let code do arithmetic and dates. 19
Good question shapes from the openjevx recipes:
| shape | decision |
|---|---|
| Find the failures in a log | yes/no per log line: keep only the lines an on-call engineer would act on 20 |
| Route a ticket queue | pick-one per ticket: which team owns it 21 |
| Check a claim against the source | yes/no: does the document say what you are about to write? 22 |
| Is this command dangerous? | yes/no about a command before running it 23 |
| Which tool should handle this? | pick-one among function names 24 |
| Several judgements in one call | one request, three questions of the three types 25 |
The recipe outputs come from model 0.5.2 on server 0.5.7, CPU, Apple M5 Pro, run on 2026-10-03, and include the wrong and unsure answers. 26 On that model, the due-date pick was wrong, and confidently: it picked the issue date (0.90). 27 On that model, the prompt-injection recipe (“does fetched text try to instruct an AI assistant?”) missed, with P(yes) 0.34. 28 29 Compute candidate dates in code; the model judges text, it does not do date arithmetic. 30
Lesson 2: pick the bar on held-out data, with controls
Section titled “Lesson 2: pick the bar on held-out data, with controls”A model can rank well while the threshold is wrong. 31
On the pitchy gate, the hosted model had AUC 0.99 but the bar was off by 0.25. 31
Context-free, the ≤ 0.20 bar still blocks 14 of 53 drafts (26%) that are not pitches. 32
Every gate needs 2-3 known-good and known-bad controls, run whenever the bar is set. 33
A plain rule written after reading the rows scored 67% in-sample, which is only an upper bound. 34
Choosing temperatures on the gate would also fit the test. 35
The OpenJevX training recipe says: hold out calibration before training, then fit one temperature per type. 36
Your inputs will score differently: copy the pattern, not the numbers, and test on your own data. 37
See Benchmarks for gate numbers and latency, and the Model card for honest limits.
Sources
Section titled “Sources”Footnotes
Section titled “Footnotes”-
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL8 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL36–37 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL38–39 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL3–8 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL8 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL6 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL12 ↩ ↩2 -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL13 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL15–17 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL21 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL23 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL24 ↩ ↩2 -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL26 ↩ ↩2 -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL27 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL29 ↩ ↩2 -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL39–40 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL41 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL50 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL84–85 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL56 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL57 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL60 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL62 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL71 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL72 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL45–46 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/04-pick-a-value.mdL48 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/08-prompt-injection.mdL1–3 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/08-prompt-injection.mdL49 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/04-pick-a-value.mdL44 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL43–44 ↩ ↩2 -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL25 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/USECASES.mdL46 ↩ -
jevresearch @ 4634ffa (4634ffa) ·
jevresearch/STATE.mdL54 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL32–33 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0002-training-run.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL50 ↩