Model card
The model in the release downloads is version 0.5.2. 1 2
Architecture
Section titled “Architecture”The base model is convaiinnovations/laya. 3
The encoder is answerdotai/ModernBERT-large (395M) plus Laya’s decision head, 421M parameters in total. 4
The training method is RLCD, Reinforcement Learning for Calibrated Decisions. 5
The model is non-autoregressive and returns probabilities in one forward pass. 6
The encoder is ModernBERT-large sized: 28 layers, hidden 1024, about 300M encoder weights. 7
The model folder
Section titled “The model folder”The server binary holds no model. 8 A model is a folder: the graph, its config and its tokenizer, and the tokenizer is optional. 9
| file | what it is |
|---|---|
openjevx.w8.onnx |
the graph, 8-bit weight-only 10 |
config.json |
name, version, temperature (choice, score, noul), max_len, head_max, special_ids, quantization, base_model, and sha256 of the onnx 11 |
tokenizer.json |
optional; the built-in ModernBERT tokenizer is used otherwise 12 |
config.json for model 0.5.2 sets max_len 512 and head_max 256. 13
The special ids are cls 50281, sep 50282, pad 50283 and mask 50284. 14
A model folder whose config.json has no "temperature" fails to load. 15 16
Temperatures must be > 0. 17
Build a folder with finetuning/export/make_model_folder.py. 18
The sha256 check
Section titled “The sha256 check”At load the server computes the sha256 of the graph file. 19
If config.json names a sha256 and it does not match, the server refuses the model. 20
GET /health reports the model version and sha256. 21
The 8-bit graph
Section titled “The 8-bit graph”The model 0.5.2 graph is openjevx.w8.onnx, 570 MiB. 22 23 24
It is QDQ weight-only: 120 per-channel int8 DequantizeLinear nodes feed fp32 MatMul. 25
With the server’s settings, ONNX Runtime fuses all 120 into MatMulNBits (8-bit, accuracy level 4 = int8 activations). 26
Why 8-bit
Section titled “Why 8-bit”The first ONNX export of the fine-tune was fp32 at 1.6GB. 27
The ship target is one small download with no material quality loss. 27
OpenJevX is a CPU-first decision model that you run in a container, and it ships one 8-bit ONNX. 28
Every release from v0.5.0 ships openjevx.w8.onnx, 8-bit weights, not the 4-bit file the ADR first chose. 29
ONNX Runtime 1.29 fuses DequantizeLinear + MatMul into MatMulNBits, so the 8-bit weights stay 8-bit in memory. 30
On ONNX Runtime 1.22 the graph rebuilt about 300 M fp32 weights on every request, which is why a 36-token request took 172 ms (model 0.5.0, Apple M5 Pro CPU). 31 32
ONNX Runtime 1.29 brought the v0.5.0 model to 22 ms short, 85 ms typical and 540 ms long (912 tokens), with the same gate numbers. 33 34
Watch out: ONNX Runtime 1.29 quantizes the activations to int8 on every call, so an answer can move slightly with what else is in the request. 35
Inputs
Section titled “Inputs”The request body is a state and a map of questions. 36
A question is parsed from its type, instructions and criteria fields. 37
The head of each encoded question is <type> question: <instructions>. 38
Each option gets a [MASK] marker token in front of it. 39
Each option is cut to 48 tokens. 40
The question head and its options share a head_max token budget: if the options leave under 16 tokens, each option is cut further, and the head keeps at least 8 tokens. 41
If the state is longer than the room left under max_len, it is truncated. 42
Many questions about one state re-encode the shared state once per question, because the question comes first in the input. 43
The ONNX graph inputs are input_ids, attention_mask, marker_pos, marker_mask and qtype. 44
The outputs are logits and act_logits. 44
Outputs
Section titled “Outputs”Every answer has type, probabilities, confidence and answer_confidence. 45
For choice, the answer’s choice is the key of the most probable option. 46
For score, the answer’s score is the sum of each level index times its probability. 47
A noul answer adds noul, the probability of yes. 48
For noul, confidence is the larger of P(true) and 1 − P(true). 49
Each answer can also carry action.act_probability, a two-way softmax over the act values. 50
The response wraps answers as {"model":"openjevx","answers":…,"usage":…}. 51
model is the string "openjevx". 52
Calibration
Section titled “Calibration”Each model folder carries its own temperatures in config.json, one per type. 53
Each logit is divided by scale, the temperature for the question’s type. 54
The scaled values are exponentiated and normalised to sum to 1, a softmax. 55
The training recipe holds out a calibration set before training and fits one temperature per type (choice, score, noul). 56
A fine-tune job writes the model folder with that run’s calibration temperatures. 57
Devices
Section titled “Devices”device is auto (default), cpu or gpu. 58
auto uses the first GPU provider that loads (CUDA, CoreML on macOS, DirectML on Windows) and whose answers match the CPU on a probe. 58
If the largest difference between GPU and CPU probe logits is above 0.25, that GPU provider is rejected. 59
With device set to gpu, startup fails with device gpu was set and no GPU provider ran the model. 60
Otherwise it logs no GPU provider ran the model, using CPU and uses the CPU. 61
Training data (model 0.5.2)
Section titled “Training data (model 0.5.2)”v0.5.2 was trained on 791,889 decisions (465,583 rows). 62
| area | decisions |
|---|---|
| Rule-checking across 38 business domains | 200,155 63 |
| Software-work roles | 183,543 64 |
| tasksource decision corpus | 146,567 65 |
| Public sets: CVE fixes from bigvul, defect detection, code search, ms_marco relevance, HDFS and BGL log alerts | 104,990 66 |
| Rule-reading drills | 75,321 67 |
| Log triage | 60,001 68 |
| Everyday basics | 12,162 69 |
| Public typed-decisions | 9,150 70 |
The run used one RTX 4090 for one full pass in 2.5 h. 71
Licence
Section titled “Licence”The licence file is the Apache License, Version 2.0. 72
The model card template for the weights says license: apache-2.0. 73
Honest limits
Section titled “Honest limits”It is not a general reasoner; on the public benchmarks it is mostly unsure (confident on 11%). 74
Known miss (model version not recorded): on real incident postmortems it is confidently wrong 18% of the time. 75
On the logs gate (900), model 0.5.2 is 7.8% confidently wrong. 76 77
On BGL (Blue Gene/L supercomputer log) alerts, v0.5.2 misses alerts it used to catch: v0.5.0 caught 58, v0.5.2 caught 30 and 32 in two runs. 78
The release gate does not cover BGL. 79
Model 0.5.0 had a saturated head: right and confidently wrong answers sat at the same raw margin, the head’s ceiling of about 4.9 logits. 80 81
A temperature cannot separate them, so recalibration could not fix it and a retrain was needed. 82
Model 0.5.2’s config.json has max_len 512, so longer inputs are now cut at 512 tokens. 13 24
The gate has no input longer than 300 tokens and can’t detect what truncation loses. 83
Long inputs and many-question batches do not run under 100 ms on the M5 Pro (model 0.5.0). 33 84 31
See Benchmarks for every number with its hardware, and Use cases for which questions work.
Sources
Section titled “Sources”Footnotes
Section titled “Footnotes”-
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL133–134 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
deploy/MODEL_VERSIONL1 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0001-base-model.mdL12 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0001-base-model.mdL14 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0001-base-model.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
scripts/publish_hf.pyL55–60 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL62 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL65 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL17–21 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL68–69 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL70–75 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL76 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL73 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL26 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL90–92 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL93–97 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL83–84 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL130–131 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/model.goL132–134 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL82–83 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL69–70 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL1 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL120–121 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-x86-cpu-latency.mdL15–17 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0003-export-and-release.mdL8 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0010-v0.5-release-and-v0.6-scope.mdL30–31 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0003-export-and-release.mdL14–15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL27 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL3 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL21–23 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0011-inference-runtime.mdL9 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0011-inference-runtime.mdL15–17 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL243–245 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/decision.goL28–31 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL64–68 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL138–139 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL141–147 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL141–146 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL153–169 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL178–181 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/12-inference-latency.mdL72–73 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/main.goL389–391 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL229–234 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL236–237 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL238–243 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
recipes/README.mdL37–39 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL244–246 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL248–252 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/decision.goL96–100 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/decision.goL96–97 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL13–16 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL200–204 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
internal/decide/decide.goL209–217 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0002-training-run.mdL15 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL257–258 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/main.goL574–576 ↩ ↩2 -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/main.goL345–355 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/main.goL381–383 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
cmd/openjevx/main.goL385–386 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL224–226 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL230 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL231 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL232 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL233 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL234 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL235 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL236 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL237 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
README.mdL239 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
LICENSEL1–2 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
scripts/publish_hf.pyL39–40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL38–39 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/11-comparison-hf-card.mdL40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL107 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL113 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL21–28 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/14-v0.5.2-benchmarks.mdL40 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL1–3 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL20–21 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
llmresults/13-v0.5.2-gate-misses.mdL22–33 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0011-inference-runtime.mdL22–23 ↩ -
openjevx @ v0.5.9 (ee2a1f4) ·
docs/adr/0011-inference-runtime.mdL27–28 ↩