Skip to content
Every statement on this page is cited and was checked against openjevx v0.5.9 (ee2a1f4) on 2026-10-04.

Model card

The model in the release downloads is version 0.5.2. 1 2

The base model is convaiinnovations/laya. 3 The encoder is answerdotai/ModernBERT-large (395M) plus Laya’s decision head, 421M parameters in total. 4 The training method is RLCD, Reinforcement Learning for Calibrated Decisions. 5 The model is non-autoregressive and returns probabilities in one forward pass. 6 The encoder is ModernBERT-large sized: 28 layers, hidden 1024, about 300M encoder weights. 7

The server binary holds no model. 8 A model is a folder: the graph, its config and its tokenizer, and the tokenizer is optional. 9

file what it is
openjevx.w8.onnx the graph, 8-bit weight-only 10
config.json name, version, temperature (choice, score, noul), max_len, head_max, special_ids, quantization, base_model, and sha256 of the onnx 11
tokenizer.json optional; the built-in ModernBERT tokenizer is used otherwise 12

config.json for model 0.5.2 sets max_len 512 and head_max 256. 13 The special ids are cls 50281, sep 50282, pad 50283 and mask 50284. 14 A model folder whose config.json has no "temperature" fails to load. 15 16 Temperatures must be > 0. 17 Build a folder with finetuning/export/make_model_folder.py. 18

At load the server computes the sha256 of the graph file. 19 If config.json names a sha256 and it does not match, the server refuses the model. 20 GET /health reports the model version and sha256. 21

The model 0.5.2 graph is openjevx.w8.onnx, 570 MiB. 22 23 24 It is QDQ weight-only: 120 per-channel int8 DequantizeLinear nodes feed fp32 MatMul. 25 With the server’s settings, ONNX Runtime fuses all 120 into MatMulNBits (8-bit, accuracy level 4 = int8 activations). 26

The first ONNX export of the fine-tune was fp32 at 1.6GB. 27 The ship target is one small download with no material quality loss. 27 OpenJevX is a CPU-first decision model that you run in a container, and it ships one 8-bit ONNX. 28 Every release from v0.5.0 ships openjevx.w8.onnx, 8-bit weights, not the 4-bit file the ADR first chose. 29 ONNX Runtime 1.29 fuses DequantizeLinear + MatMul into MatMulNBits, so the 8-bit weights stay 8-bit in memory. 30 On ONNX Runtime 1.22 the graph rebuilt about 300 M fp32 weights on every request, which is why a 36-token request took 172 ms (model 0.5.0, Apple M5 Pro CPU). 31 32 ONNX Runtime 1.29 brought the v0.5.0 model to 22 ms short, 85 ms typical and 540 ms long (912 tokens), with the same gate numbers. 33 34

Watch out: ONNX Runtime 1.29 quantizes the activations to int8 on every call, so an answer can move slightly with what else is in the request. 35

The request body is a state and a map of questions. 36 A question is parsed from its type, instructions and criteria fields. 37 The head of each encoded question is <type> question: <instructions>. 38 Each option gets a [MASK] marker token in front of it. 39 Each option is cut to 48 tokens. 40 The question head and its options share a head_max token budget: if the options leave under 16 tokens, each option is cut further, and the head keeps at least 8 tokens. 41 If the state is longer than the room left under max_len, it is truncated. 42 Many questions about one state re-encode the shared state once per question, because the question comes first in the input. 43

The ONNX graph inputs are input_ids, attention_mask, marker_pos, marker_mask and qtype. 44 The outputs are logits and act_logits. 44

Every answer has type, probabilities, confidence and answer_confidence. 45 For choice, the answer’s choice is the key of the most probable option. 46 For score, the answer’s score is the sum of each level index times its probability. 47 A noul answer adds noul, the probability of yes. 48 For noul, confidence is the larger of P(true) and 1 − P(true). 49 Each answer can also carry action.act_probability, a two-way softmax over the act values. 50 The response wraps answers as {"model":"openjevx","answers":…,"usage":…}. 51 model is the string "openjevx". 52

Each model folder carries its own temperatures in config.json, one per type. 53 Each logit is divided by scale, the temperature for the question’s type. 54 The scaled values are exponentiated and normalised to sum to 1, a softmax. 55 The training recipe holds out a calibration set before training and fits one temperature per type (choice, score, noul). 56 A fine-tune job writes the model folder with that run’s calibration temperatures. 57

device is auto (default), cpu or gpu. 58 auto uses the first GPU provider that loads (CUDA, CoreML on macOS, DirectML on Windows) and whose answers match the CPU on a probe. 58 If the largest difference between GPU and CPU probe logits is above 0.25, that GPU provider is rejected. 59 With device set to gpu, startup fails with device gpu was set and no GPU provider ran the model. 60 Otherwise it logs no GPU provider ran the model, using CPU and uses the CPU. 61

v0.5.2 was trained on 791,889 decisions (465,583 rows). 62

area decisions
Rule-checking across 38 business domains 200,155 63
Software-work roles 183,543 64
tasksource decision corpus 146,567 65
Public sets: CVE fixes from bigvul, defect detection, code search, ms_marco relevance, HDFS and BGL log alerts 104,990 66
Rule-reading drills 75,321 67
Log triage 60,001 68
Everyday basics 12,162 69
Public typed-decisions 9,150 70

The run used one RTX 4090 for one full pass in 2.5 h. 71

The licence file is the Apache License, Version 2.0. 72 The model card template for the weights says license: apache-2.0. 73

It is not a general reasoner; on the public benchmarks it is mostly unsure (confident on 11%). 74 Known miss (model version not recorded): on real incident postmortems it is confidently wrong 18% of the time. 75 On the logs gate (900), model 0.5.2 is 7.8% confidently wrong. 76 77 On BGL (Blue Gene/L supercomputer log) alerts, v0.5.2 misses alerts it used to catch: v0.5.0 caught 58, v0.5.2 caught 30 and 32 in two runs. 78 The release gate does not cover BGL. 79 Model 0.5.0 had a saturated head: right and confidently wrong answers sat at the same raw margin, the head’s ceiling of about 4.9 logits. 80 81 A temperature cannot separate them, so recalibration could not fix it and a retrain was needed. 82 Model 0.5.2’s config.json has max_len 512, so longer inputs are now cut at 512 tokens. 13 24 The gate has no input longer than 300 tokens and can’t detect what truncation loses. 83 Long inputs and many-question batches do not run under 100 ms on the M5 Pro (model 0.5.0). 33 84 31

See Benchmarks for every number with its hardware, and Use cases for which questions work.

  1. openjevx @ v0.5.9 (ee2a1f4) · README.md L133–134 ↩

  2. openjevx @ v0.5.9 (ee2a1f4) · deploy/MODEL_VERSION L1 ↩

  3. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0001-base-model.md L12 ↩

  4. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0001-base-model.md L14 ↩

  5. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0001-base-model.md L15 ↩

  6. openjevx @ v0.5.9 (ee2a1f4) · scripts/publish_hf.py L55–60 ↩

  7. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L62 ↩

  8. openjevx @ v0.5.9 (ee2a1f4) · README.md L65 ↩

  9. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L17–21 ↩

  10. openjevx @ v0.5.9 (ee2a1f4) · README.md L68–69 ↩

  11. openjevx @ v0.5.9 (ee2a1f4) · README.md L70–75 ↩

  12. openjevx @ v0.5.9 (ee2a1f4) · README.md L76 ↩

  13. openjevx @ v0.5.9 (ee2a1f4) · README.md L70–72 ↩ ↩2

  14. openjevx @ v0.5.9 (ee2a1f4) · README.md L73 ↩

  15. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L26 ↩

  16. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L90–92 ↩

  17. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L93–97 ↩

  18. openjevx @ v0.5.9 (ee2a1f4) · README.md L83–84 ↩

  19. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L130–131 ↩

  20. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/model.go L132–134 ↩

  21. openjevx @ v0.5.9 (ee2a1f4) · README.md L82–83 ↩

  22. openjevx @ v0.5.9 (ee2a1f4) · README.md L69–70 ↩

  23. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L1 ↩

  24. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L120–121 ↩ ↩2

  25. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L15 ↩

  26. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-x86-cpu-latency.md L15–17 ↩

  27. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0003-export-and-release.md L8 ↩ ↩2

  28. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0010-v0.5-release-and-v0.6-scope.md L30–31 ↩

  29. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0003-export-and-release.md L14–15 ↩

  30. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L27 ↩

  31. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L3 ↩ ↩2

  32. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L21–23 ↩

  33. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0011-inference-runtime.md L9 ↩ ↩2

  34. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0011-inference-runtime.md L15–17 ↩

  35. openjevx @ v0.5.9 (ee2a1f4) · README.md L243–245 ↩

  36. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/decision.go L28–31 ↩

  37. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L64–68 ↩

  38. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L138–139 ↩

  39. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L141–147 ↩

  40. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L141–146 ↩

  41. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L153–169 ↩

  42. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L178–181 ↩

  43. openjevx @ v0.5.9 (ee2a1f4) · llmresults/12-inference-latency.md L72–73 ↩

  44. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/main.go L389–391 ↩ ↩2

  45. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L229–234 ↩

  46. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L236–237 ↩

  47. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L238–243 ↩

  48. openjevx @ v0.5.9 (ee2a1f4) · recipes/README.md L37–39 ↩

  49. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L244–246 ↩

  50. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L248–252 ↩

  51. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/decision.go L96–100 ↩

  52. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/decision.go L96–97 ↩

  53. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L13–16 ↩

  54. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L200–204 ↩

  55. openjevx @ v0.5.9 (ee2a1f4) · internal/decide/decide.go L209–217 ↩

  56. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0002-training-run.md L15 ↩

  57. openjevx @ v0.5.9 (ee2a1f4) · README.md L257–258 ↩

  58. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/main.go L574–576 ↩ ↩2

  59. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/main.go L345–355 ↩

  60. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/main.go L381–383 ↩

  61. openjevx @ v0.5.9 (ee2a1f4) · cmd/openjevx/main.go L385–386 ↩

  62. openjevx @ v0.5.9 (ee2a1f4) · README.md L224–226 ↩

  63. openjevx @ v0.5.9 (ee2a1f4) · README.md L230 ↩

  64. openjevx @ v0.5.9 (ee2a1f4) · README.md L231 ↩

  65. openjevx @ v0.5.9 (ee2a1f4) · README.md L232 ↩

  66. openjevx @ v0.5.9 (ee2a1f4) · README.md L233 ↩

  67. openjevx @ v0.5.9 (ee2a1f4) · README.md L234 ↩

  68. openjevx @ v0.5.9 (ee2a1f4) · README.md L235 ↩

  69. openjevx @ v0.5.9 (ee2a1f4) · README.md L236 ↩

  70. openjevx @ v0.5.9 (ee2a1f4) · README.md L237 ↩

  71. openjevx @ v0.5.9 (ee2a1f4) · README.md L239 ↩

  72. openjevx @ v0.5.9 (ee2a1f4) · LICENSE L1–2 ↩

  73. openjevx @ v0.5.9 (ee2a1f4) · scripts/publish_hf.py L39–40 ↩

  74. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L38–39 ↩

  75. openjevx @ v0.5.9 (ee2a1f4) · llmresults/11-comparison-hf-card.md L40 ↩

  76. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L107 ↩

  77. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L113 ↩

  78. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L21–28 ↩

  79. openjevx @ v0.5.9 (ee2a1f4) · llmresults/14-v0.5.2-benchmarks.md L40 ↩

  80. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L1–3 ↩

  81. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L20–21 ↩

  82. openjevx @ v0.5.9 (ee2a1f4) · llmresults/13-v0.5.2-gate-misses.md L22–33 ↩

  83. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0011-inference-runtime.md L22–23 ↩

  84. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0011-inference-runtime.md L27–28 ↩