Skip to content
Every statement on this page is cited and was checked against jev-cloud origin/main (10a744c), openjevx v0.5.9 (ee2a1f4) on 2026-10-04.

Fine-tune

There are two paths. Path A runs the OpenJevX fine-tuning pipeline yourself. Path B runs a SageMaker training job in your own AWS account with jev-train. Both start from the same CSV.

Two limits apply to both paths today. Read them before you rent a GPU.

Warning: where training starts. In the openjevx pipeline, task all does not yet start from the checkpoint, so every run trains from the base model. 1 In the SageMaker training image, the default base_model is convaiinnovations/laya. 2 3

Path A: each run saves a trainable checkpoint next to the model, but task all does not yet start from it, so today every run trains from the base model on the whole mix. 1

The SageMaker training image takes ONE customer’s decisions.csv and builds shards from those rows only. 4

Path B: the openjevx code runs unchanged from the pinned tag v0.5.0. 5

SageMaker job minimum: a CSV needs ≥ 50 train rows for smoke and ≥ 34 for full. 6 openjevx’s own examples/decisions.csv has 53 train rows since 2026-10-03; the 39-row version before it (24 train rows) failed both. 7 By the same ×2 / ×3 repeats that is 106 (smoke) and 159 (full) items, above 100, computed from validate’s split and not yet measured by a dry run. 8 jev-cloud ships client/examples/decisions.csv; its test asserts it validates, has enough train rows for both smoke and full, and has held-out gate rows. 9

The old 39-row example’s dry run failed: 48 items < 100. 10

One row is one decision: put the fact the rule needs in the state, ask the question, give the right answer. 11

Add near-miss rows around each threshold (499 / 500 / 501) so the model learns the line, not the topic. 12

The file is UTF-8, comma-separated, with a header row. 13

Column Required What it holds
state yes the facts, as a JSON object or plain text, e.g. {"order_total_rs": 501} 14
question yes The instruction, for example “Does the order get the discount? Orders over Rs 500 get 10% off.” 15
type no (default noul) noul (yes/no), choice or score. 16
options for choice and score the option names, as a|b|c or a=what a means|b=... 17
answer yes noul: yes/no/true/false/1/0; choice: an option name; score: a level name or index (0 = lowest). 18
split no train, test or gate; empty means 90% train / 10% gate, fixed per row. 19
source no (default your-data) A tag for where the row came from. 20

options is a|b|c or a=what a means|b=...; score levels go lowest first; for noul, leave it empty or give true=...|false=.... 17

An empty split defaults to 90% train / 10% gate, fixed by a hash of state+question. 21

jev-train also accepts yes/no, yesno, bool and boolean as noul. 22 23

JSON inside a CSV cell: wrap the cell in double quotes and double the quotes inside, as in "{""day"": ""Friday""}". 24

A state that starts with { must be valid JSON and a non-empty object; anything else is plain text. 25

A cell that contains a comma must be quoted, or the row has more cells than the header. 26

  • Aim for 20+ rows per question. 27
  • Keep no single answer above 80%. 27
  • validate reports both as warnings: fewer than 20 rows for a question, or one answer above 80%. 28
  • Two rows with the same state and question but different answers are an error: a fact that decides the answer is missing from the state. 29
  • Never put the answer in the state ("eligible": true). 30
  • Keep a few of the hardest rows as split=gate, so you can measure the result. 31

Training rows have the shape {state, questions:{id:{type: noul|choice|score, instructions, criteria}}, gold}. 32

New domain data must already be state + questions + gold.probabilities; hard labels alone are not an RLCD target. 33

The gold field holds full probability distributions. 34

Terminal window
python finetuning/dataprep/import_csv.py yours.csv --name mydata --add-to-config

The importer writes <data>/train/NAME_train.jsonl, <data>/eval/NAME_eval.jsonl and <data>/gate/NAME_gate.jsonl. 35

It refuses to write if any row is bad, unless --skip-bad. 36

--dry-run checks without writing. 37

--add-to-config adds your train file to the training mix (repeated 3 times; change with --repeat), your gate file to the gate, and both to the leakage check. 38

Settings live in ~/.config/openjevx/config.json, created from finetuning/config.example.json; override the path with $OPENJEVX_FT_CONFIG. 39

Data lives in ~/openjevx/data/; override it with $OPENJEVX_DATA. 40

The data folder holds raw/, incoming/, train/, eval/, gate/ and work/{leak,shards,runs,gate,quality,samples}. 41

One runner, ft.py, drives every stage, and every step fails loudly. 42 43

Command What it does
ft.py dataprep Generates the rule-labelled sets into <data>/{train,eval,gate}. 44
ft.py validate Adapter dry-parse plus leakage check; fails if a gate question asks about a state that is in training. 45
ft.py package [--smoke] Builds the shard in <data>/work/shards/<version>[-smoke]/ and runs every row through the trainer’s build_item on CPU. 46
ft.py train [--smoke] Runs the shard on config.provider and gets the 8-bit ONNX back. 47
ft.py gate MODEL Serves MODEL locally, scores it, and exits 1 if it misses the config thresholds. 48
ft.py all Every step in order, with a smoke run before the full run. 49

The same stages run from the Taskfile in finetuning/: task all, task smoke, task train, task gate -- <model folder>. 50

Terminal window
cd finetuning
task all

task all runs dataprep → validate → smoke → full train → gate, stopping at the first failure. 51

Every run is one job, end to end: data ready → validate (leakage check) → smoke run → full run → gate. 52

Freeze the data before renting a GPU; a model ships only when the gate passes, not when training ends. 53

config.example.json sets the provider to vast with GPU RTX_4090. 54

The laptop uploads the shard to the private bucket openjevx-train and a job.env holding only signed links and the kill token. 55

The box uses pytorch/pytorch:2.6.0-cuda12.4-cudnn9-runtime and installs the packages in finetuning/train/requirements-box.txt from PyPI. 56

The box trains, calibrates, exports ONNX on CUDA, quantizes to 8-bit, and fails if the 8-bit ONNX is over 750 MB. 57

It uploads openjevx.w8.onnx, checkpoint.tar.gz, eval_report.json, job.log and status.json to runs/<run>/ through signed PUT links. 58

The laptop only waits on R2: gpu/vast.py polls runs/<run>/status.json and downloads the results to ~/openjevx/data/work/runs/<run>/out/. 59

A run survives the laptop sleeping or the agent session ending; results wait in R2. 60

task cloud-setup creates or refreshes the private R2 bucket and the self-destroy endpoint, and checks it. 61

The box clones a pushed commit, so commit and push any change under finetuning/ before training. 62

From a fork, origin must be a public repo you can push to, because the box clones it over HTTPS with no credentials. 63

Another GPU provider is one file, finetuning/gpu/<name>.py, that prints RUN_DIR=<dir> and leaves the model folder in <dir>/out/model/. 64

On your own CUDA box, run the steps in finetuning/ft.py and finetuning/train/run_job.sh directly. 65

The training box must use onnxruntime-gpu==1.22.0; version 1.30 needs CUDA 13 and silently falls back to CPU on this image. 66

Terminal window
python3 ft.py gate ~/openjevx/data/work/runs/<run>/out/model
Threshold (config gate) Default
min_basics_confident_right 0.9 67
max_basics_confident_wrong 0.02 68
min_jevx13_correct 12 69

The gate also scores the logs gate and the held-out test sets, and nothing may regress against the previous release. 70

Right & confident means right and sure enough to act on: yes ≥ 0.8, no ≤ 0.2, a choice or score ≥ 0.6. 71

Confidently wrong means sure and wrong; keep it near zero. 72

The full report is saved to ~/openjevx/data/work/gate/<model>.json. 73

Honest limit: a small gate file is noisy. A CSV with ~15 held-out rows gives a very noisy accuracy, so add split=test rows. 74

jev-train validates your CSV on your machine, uploads it to an S3 bucket in your account, starts a SageMaker training job in your account, and watches it to the end. 75

validate uses exactly the rules of the OpenJevX importer the trainer uses. 76

It needs Node 20 or newer. 77

Credentials come from the AWS SDK default chain only, and jev-train never asks for keys and never prints them. 78

Terminal window
jev-train validate decisions.csv
jev-train upload decisions.csv --bucket my-jev-models [--job NAME]
jev-train start --bucket my-jev-models --role-arn arn:aws:iam::123456789012:role/JevTraining \
--image 123456789012.dkr.ecr.us-east-1.amazonaws.com/jev-train:v0.5.0 \
[--job NAME] [--instance ml.g4dn.xlarge] [--spot] [--max-hours 3] [--smoke]
jev-train status jev-20261002-101500
jev-train watch jev-20261002-101500 [--interval 30]
jev-train run decisions.csv --bucket ... --role-arn ... --image ... [start options]
Command What it does
validate <csv> Local only; prints every bad row with its line number and a fix, then rows and answers per question, with warnings below 20 rows or above 80% for one answer. 28
upload <csv> --bucket B Validates first, then uploads to s3://B/training/<job>/input/decisions.csv. 79
start Calls CreateTrainingJob. 80
status <job> Prints the status, the secondary status and the billable seconds. 81
watch <job> Polls every --interval seconds (30 by default); Ctrl-C stops watching and the job keeps running. 82
run <csv> validate, upload, start, watch. 83

--stack NAME reads the defaults from a CloudFormation stack’s outputs: ModelsBucket, TrainingRoleArn and TrainingImage. 84

Any flag you pass overrides the stack’s value. 85

The job name defaults to jev-YYYYMMDD-HHMMSS (UTC). 79

S3 path What
s3://B/training/<job>/input/decisions.csv the customer CSV, uploaded by the jev client after local validation 86 87
s3://B/training/<job>/output/model.tar.gz SageMaker’s own output. 86 88
s3://B/models/current/ What the server loads. 89
s3://B/models/<version>/ Every trained or delivered version, immutable once written. 90 91

The job runs on --instance (default ml.g4dn.xlarge), 1 instance, 50 GB volume. 92

The hyperparameter shard is smoke with --smoke, otherwise full. 93

The default base model is baked into the image, so the job runs with HF_HUB_OFFLINE=1; any other base_model repo is downloaded at train time and needs network access. 3

Part Path Notes
command docker run <image> train Any other argument prints usage and exits 2. 94
input /opt/ml/input/data/train/decisions.csv Channel train; if decisions.csv is missing, the only *.csv in the channel is used. 95
hyperparameters /opt/ml/input/config/hyperparameters.json All values are strings. 96
model output /opt/ml/model/ → S3OutputPath/<job>/output/model.tar.gz Files at the root of the tar. 97
checkpoint /opt/ml/output/data/checkpoint/ The trainable checkpoint for a later re-fine-tune, kept out of the model tar. 98
failure /opt/ml/output/failure + non-zero exit SageMaker shows it as FailureReason. 99
logs stdout → CloudWatch CSV contents are never logged. 100

model.tar.gz holds openjevx.w8.onnx, config.json, tokenizer.json, eval_report.json and manifest.json, all at the root. 101

Hyperparameter Default Meaning
shard smoke smoke: 1 epoch, micro-batch 2, accuracy gate 0.0; full: 4 epochs, micro 8 × accum 8, gate 0.55. 102
base_model convaiinnovations/laya A Hugging Face repo in laya layout. 3
max_w8_mb 750 Fail if the 8-bit ONNX is larger, in MiB. 103
dry_run 0 1 stops before GPU training. 104

A dry run needs no GPU, so a cheap CPU instance (ml.m5.large) can check the data first. 105

The job needs at least 34 rows in the train split for a full run, or 50 for --smoke. 106

The reason: train_job.py raises below 100 trainable items, and the smoke shard holds your train rows ×2 and the full shard ×3. 107

A full run also needs held-out rows (split test or gate). 108

Warning: the container has not run on a GPU yet. The runtime figures in its README are estimates, not measurements. 109

On a T4, do not set AMP=bf16. 110

A model is one folder: openjevx.w8.onnx (8-bit weight-only graph), config.json and tokenizer.json. 111

The job produces the model folder <run>/out/model/ with that run’s calibration temperatures; set "model" in openjevx.json to it. 112

{ "listen": "127.0.0.1:21118", "device": "auto", "model": "/path/to/my-model" }

GET /health reports version and sha256, so you can check which model is live. 113

On S3, promote means copying the 3 files of models/<version>/ into models/current/; a server ETag reload or a rolling restart picks it up. 114

Rollback means promoting the previous version again. 115

Deploy under a new model name (mymodel-v2), so no cached answer from the old model is reused. 116

Problems on the way: see Troubleshooting. Terms: see Glossary.

  1. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L299–300 ↩ ↩2

  2. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L1 ↩

  3. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L47 ↩ ↩2 ↩3

  4. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L3–5 ↩

  5. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L9–10 ↩

  6. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L65–66 ↩

  7. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L66–68 ↩

  8. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L107–108 ↩

  9. jev-cloud @ origin/main (10a744c) · client/test/example.test.js L8–15 ↩

  10. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L100–102 ↩

  11. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L10 ↩

  12. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L11 ↩

  13. openjevx @ v0.5.9 (ee2a1f4) · finetuning/dataprep/import_csv.py L6 ↩

  14. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L15 ↩

  15. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L16 ↩

  16. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L17 ↩

  17. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L18 ↩ ↩2

  18. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L19 ↩

  19. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L20 ↩

  20. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L21 ↩

  21. openjevx @ v0.5.9 (ee2a1f4) · finetuning/dataprep/import_csv.py L14 ↩

  22. jev-cloud @ origin/main (10a744c) · client/README.md L1 ↩

  23. jev-cloud @ origin/main (10a744c) · client/README.md L79 ↩

  24. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L23–24 ↩

  25. jev-cloud @ origin/main (10a744c) · client/README.md L87–88 ↩

  26. jev-cloud @ origin/main (10a744c) · client/README.md L89 ↩

  27. openjevx @ v0.5.9 (ee2a1f4) · finetuning/examples/README.md L26 ↩ ↩2

  28. jev-cloud @ origin/main (10a744c) · client/README.md L42 ↩ ↩2

  29. jev-cloud @ origin/main (10a744c) · client/README.md L90–91 ↩

  30. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L122 ↩

  31. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L121 ↩

  32. openjevx @ v0.5.9 (ee2a1f4) · README.md L250 ↩

  33. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0001-base-model.md L22 ↩

  34. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0002-training-run.md L14 ↩

  35. openjevx @ v0.5.9 (ee2a1f4) · finetuning/dataprep/import_csv.py L17 ↩

  36. openjevx @ v0.5.9 (ee2a1f4) · finetuning/dataprep/import_csv.py L18 ↩

  37. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L144 ↩

  38. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L144–146 ↩

  39. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L7–8 ↩

  40. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L9 ↩

  41. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L9–10 ↩

  42. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L2 ↩

  43. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L8–16 ↩

  44. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L8 ↩

  45. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L9–10 ↩

  46. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L11–12 ↩

  47. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L13 ↩

  48. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L14–15 ↩

  49. openjevx @ v0.5.9 (ee2a1f4) · finetuning/ft.py L16 ↩

  50. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L6 ↩

  51. openjevx @ v0.5.9 (ee2a1f4) · finetuning/Taskfile.yml L28–30 ↩

  52. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L20 ↩

  53. openjevx @ v0.5.9 (ee2a1f4) · AGENTS.md L20–21 ↩

  54. openjevx @ v0.5.9 (ee2a1f4) · finetuning/config.example.json L3–6 ↩

  55. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L21–24 ↩

  56. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L25–27 ↩

  57. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L31–32 ↩

  58. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L32–34 ↩

  59. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L39–40 ↩

  60. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0009-one-job-gpu-run-via-r2.md L60 ↩

  61. openjevx @ v0.5.9 (ee2a1f4) · finetuning/Taskfile.yml L7–9 ↩

  62. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L186–187 ↩

  63. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L189–191 ↩

  64. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L198–199 ↩

  65. openjevx @ v0.5.9 (ee2a1f4) · README.md L255–256 ↩

  66. openjevx @ v0.5.9 (ee2a1f4) · docs/adr/0002-training-run.md L18 ↩

  67. openjevx @ v0.5.9 (ee2a1f4) · finetuning/config.example.json L140 ↩

  68. openjevx @ v0.5.9 (ee2a1f4) · finetuning/config.example.json L141 ↩

  69. openjevx @ v0.5.9 (ee2a1f4) · finetuning/config.example.json L142 ↩

  70. openjevx @ v0.5.9 (ee2a1f4) · docs/RELEASE_PROCESS.md L77–79 ↩

  71. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L217 ↩

  72. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L219 ↩

  73. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L223 ↩

  74. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L151–152 ↩

  75. jev-cloud @ origin/main (10a744c) · client/README.md L5–9 ↩

  76. jev-cloud @ origin/main (10a744c) · client/README.md L5–6 ↩

  77. jev-cloud @ origin/main (10a744c) · client/README.md L15 ↩

  78. jev-cloud @ origin/main (10a744c) · client/README.md L23–24 ↩

  79. jev-cloud @ origin/main (10a744c) · client/README.md L43 ↩ ↩2

  80. jev-cloud @ origin/main (10a744c) · client/README.md L44 ↩

  81. jev-cloud @ origin/main (10a744c) · client/README.md L45 ↩

  82. jev-cloud @ origin/main (10a744c) · client/README.md L46 ↩

  83. jev-cloud @ origin/main (10a744c) · client/README.md L47 ↩

  84. jev-cloud @ origin/main (10a744c) · client/README.md L49–50 ↩

  85. jev-cloud @ origin/main (10a744c) · client/README.md L50 ↩

  86. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L5 ↩ ↩2

  87. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L16–17 ↩

  88. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L16–18 ↩

  89. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L8 ↩

  90. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L6–7 ↩

  91. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L12 ↩

  92. jev-cloud @ origin/main (10a744c) · client/README.md L58 ↩

  93. jev-cloud @ origin/main (10a744c) · client/README.md L60 ↩

  94. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L20 ↩

  95. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L21 ↩

  96. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L22 ↩

  97. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L24 ↩

  98. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L25 ↩

  99. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L26 ↩

  100. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L27 ↩

  101. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L29–37 ↩

  102. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L46 ↩

  103. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L48 ↩

  104. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L49 ↩

  105. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L86–87 ↩

  106. jev-cloud @ origin/main (10a744c) · client/README.md L94–95 ↩

  107. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L64–66 ↩

  108. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L69 ↩

  109. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L123 ↩

  110. jev-cloud @ origin/main (10a744c) · train/sagemaker/README.md L117 ↩

  111. openjevx @ v0.5.9 (ee2a1f4) · docs/RELEASE_PROCESS.md L61–64 ↩

  112. openjevx @ v0.5.9 (ee2a1f4) · README.md L257–259 ↩

  113. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L264–265 ↩

  114. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L20–21 ↩

  115. jev-cloud @ origin/main (10a744c) · docs/model-s3-contract.md L21 ↩

  116. openjevx @ v0.5.9 (ee2a1f4) · docs/book/train-your-own-jev.md L302 ↩