Fine-tune it on your decisions
The shipped model does not know your rules: that your week is four days, that a discount starts above Rs 500, that repo billing belongs to the payments team. Write those decisions in a spreadsheet, and one command trains a model that does. Each step below says what you should see.
0. What you get and what it costs
You get your own openjevx.w8.onnx (8-bit, at most 750 MB) plus a trainable checkpoint. The model runs on your machine; your decisions never go to a hosted model.
Training runs on a rented GPU (an RTX 4090 on vast.ai, about $0.40–0.60 an hour). One run costs about $1–3. For reference, the v0.5.0 full run took about 2.4 hours on 734k decisions and cost about $1.10; the smoke run before it costs cents. The box deletes itself when it is done.
| figure | what it is |
|---|---|
| ~$1–3 | one fine-tune run |
| ~2.4 h | v0.5.0 full run on a 4090 |
| ≤ 750 MB | your 8-bit model |
1. Install
You need Go, Task, uv and Python 3.11+ on your machine.
Get the public fine-tune kit: scripts, the trainable checkpoint and the ONNX model (v0.4.0, 2.2 GB). Verify the checksum before you unpack it:
curl -L -O https://pub-8da821f06ff747cda688f8267ed2aa96.r2.dev/openjevx-finetune-v0.4.0.tar.gz
curl -L -O https://pub-8da821f06ff747cda688f8267ed2aa96.r2.dev/openjevx-finetune-v0.4.0.tar.gz.sha256
shasum -a 256 -c openjevx-finetune-v0.4.0.tar.gz.sha256
tar -xzf openjevx-finetune-v0.4.0.tar.gz
cd openjevx-finetune-v0.4.0
go version && task --version && uv --version && python3 --version
You should see openjevx-finetune-v0.4.0.tar.gz: OK. The kit unpacks to model.safetensors (the trainable checkpoint), openjevx.onnx and openjevx.w8.onnx, encoder/, tokenizer/, and scripts/ (the trainer train_openjevx.py, quantize_w8.py, export_onnx_gpu.py, and the vast/ job scripts).
The one-command pipeline used in steps 3 to 6 (import_csv.py, the finetuning/ Taskfile with task cloud-setup, task gate and task all) is not inside the kit. Ask us for it: [email protected]. Commands below assume you have it in a folder named openjevx.
For the GPU step you also need:
- the
vastaiCLI and a vast.ai API key with a few dollars of credit (pip install vastai && vastai set api-key YOUR_KEY_HERE); - a Cloudflare account for R2 storage: the GPU box downloads your data from R2 and uploads the model back there.
cd finetuning && task cloud-setupcreates the private bucket and the self-destroy endpoint, then checks it; you should see the endpoint answer.
Optional: the jevx client (for the 13-fundamentals check in the gate and for using the model afterwards; ask us for it) and Docker. Steps 2 to 4 need no GPU and no cloud account.
2. What goes in your CSV
A row is one decision you want the model to make.
state= the facts at the moment of deciding: what your code, log or ticket knows.question= the decision, phrased as you would ask a colleague.answer= what the right decision was, from your rule or from what really happened. Never a guess.
Start from the example. It covers a 4-day week, 3 working hours a day, a late parcel, a discount over Rs 500, team routing by repo and ticket urgency: Download decisions.csv (39 rows).
The columns
| column | required? | what to put | good example | common mistake |
|---|---|---|---|---|
state | yes | the facts, as a JSON object or plain text; include the threshold if it varies | {"hours_logged_today": 3.1} | just true, or leaving out the fact the rule needs |
question | yes | one decision, as you'd ask a colleague | Is the parcel late? | two decisions in one: "is it late and who handles it" |
type | no (noul) | noul = yes/no, choice = pick one, score = rating | choice | using choice for yes/no |
options | choice, score | the keys the model picks between, |-separated; optional short meaning after =; score levels lowest first | payments=billing repos|web=frontend repos | score levels out of order (high|low|medium) |
answer | yes | yes/no (also true/false, 1/0); one option key; a level name or index (0 = lowest) | yes, payments, high | an answer that is not one of the options |
split | no | train, test or gate; empty = 90% train, 10% gate, fixed per row | gate | no gate rows, so you can't measure the result |
source | no | a tag for where the row came from | shop | — |
JSON in a CSV cell: wrap the cell in double quotes and double the quotes inside ("{""day"": ""Friday""}"). Any spreadsheet does this when you save as CSV.
One example per type: the CSV row, and what the model receives
Each row becomes one request to the server, POST /v1/systemone, with your state and question. Training teaches the model to give your answer to that request.
Yes/no (noul)
CSV row:
state,question,type,options,answer
"{""order_total_rs"": 501}","Does the order get the discount? Orders over Rs 500 get 10% off.",noul,,yes
Request the model receives:
{
"state": {"order_total_rs": 501},
"questions": {"q1": {"type": "noul",
"instructions": "Does the order get the discount? Orders over Rs 500 get 10% off."}}
}
Pick one (choice)
CSV row:
state,question,type,options,answer
"{""repo"": ""billing"", ""title"": ""Invoice PDF missing tax line""}",Which team owns this?,choice,payments=billing and payments-core repos|platform=infra-terraform and api-gateway repos|web=web-frontend and admin-console repos,payments
Request the model receives:
{
"state": {"repo": "billing", "title": "Invoice PDF missing tax line"},
"questions": {"q1": {"type": "choice", "instructions": "Which team owns this?",
"criteria": {"payments": "billing and payments-core repos",
"platform": "infra-terraform and api-gateway repos",
"web": "web-frontend and admin-console repos"}}}
}
Rating (score)
CSV row:
state,question,type,options,answer
"{""customers_affected"": 40, ""payments_failing"": false, ""workaround"": false}",How urgent is this ticket?,score,low|medium|high|critical,high
Request the model receives:
{
"state": {"customers_affected": 40, "payments_failing": false, "workaround": false},
"questions": {"q1": {"type": "score", "instructions": "How urgent is this ticket?",
"criteria": ["low", "medium", "high", "critical"]}}
}
Where the facts come from
Put in the state exactly what your system will send the model when it asks for real.
| source | state |
|---|---|
| a log alert | {"line": "OOMKilled payments-api", "env": "prod", "errors_last_10m": 42} |
| a support ticket | {"text": "Charged twice for order 8812", "customer_tier": "gold"} |
| a database record | {"promised_by": "2026-10-03", "delivered_on": "2026-10-04"} |
| a pull request | {"files_changed": 14, "touches": ["migrations/"], "summary": "adds a column"} |
| a form | {"leave_days_requested": 6, "leave_days_left": 4} |
Do and don't
| do | don't |
|---|---|
| near-miss pairs at every threshold: 499 → no, 500 → no, 501 → yes; Thursday → yes, Friday → no | put the answer inside the state ("eligible": true) |
| balance the answers: no single answer above 80% of a question's rows | train on another model's guesses: labels come from the rule or from reality |
| 20+ rows per question | give the same state and question two different answers (a fact is missing) |
keep a few of the hardest rows as split=gate | copy one wording everywhere; vary how the question and facts are phrased |
How many rows
Start with 200–1,000 rows. Each rule becomes many rows: both sides of the threshold, different values, different wordings. The import step warns you when a question has fewer than 20 rows or one answer takes more than 80% of them. The example CSV is deliberately small (39 rows) to show the shape; it gets those warnings too.
3. Import and check
python3 finetuning/dataprep/import_csv.py yours.csv --name mydata --add-to-config
You should see the count of good and bad rows, then one line per question with its answers:
yours.csv: 39 good rows, 0 bad
question rows labels
Does the order get the discount? Orders over Rs 500 get 10% 6 false=3, true=3
Which team owns this? 7 platform=3, payments=2, web=2
warning: 'Which team owns this?': only 7 rows; aim for 20+ with near-miss pairs around the rule
adapter: every row accepted
wrote 24 train rows -> ~/openjevx/data/train/mydata_train.jsonl
wrote 15 gate rows -> ~/openjevx/data/gate/mydata_gate.jsonl
updated ~/.config/openjevx/config.json (backup: config.json.bak-…)
A bad row is reported with its line number and a fix (line 7: answer 'maybe' is not yes/no). Nothing is written until every row is good, unless you pass --skip-bad. Use --dry-run to check without writing.
--add-to-config backs up ~/.config/openjevx/config.json, then adds your train file to the training mix (repeated 3 times; change with --repeat), your gate file to the gate, and both to the leakage check. Re-running it does not add duplicates.
4. Test the current model on your gate first
Before paying for a GPU, see how the shipped model already does on your rules. You need the v0.5.0 model folder (openjevx.w8.onnx, config.json, tokenizer.json, 467 MB). It is not on a public download link yet: ask us for it at [email protected]. Unpack it next to the pipeline, then point the gate at it:
tar -xzf openjevx-model-0.5.0.tar.gz # -> model/ (openjevx.w8.onnx, config.json, tokenizer.json)
cd finetuning
task gate -- ../model
This builds and serves the model on a free local port and scores every gate file, one line each:
gate/mydata_gate.jsonl n= 15 accuracy 60.0% right&confident 40.0% confidently WRONG 13.3%
That line is your baseline. If your rules already score well, you may not need to train at all.
5. Train
cd finetuning
task all
It runs these stages and stops at the first failure:
- dataprep generates the built-in rule-labelled sets (basics, conditions, IT work, logs).
- leakage check removes any training question that also appears in a test or gate file.
- trainer check on CPU builds the shard and runs rows through the real trainer, so a data bug fails here, for free, not on a rented GPU.
- smoke run: about 10 minutes on a GPU to prove the whole path works.
- full run: the real training, a few hours.
- gate on the new model.
The GPU box downloads the shard from your private R2 bucket, trains, exports and quantizes to 8-bit, uploads the model, checkpoint and log back to R2, then destroys itself. A timer destroys it at the deadline no matter what. Your laptop only waits on R2, so it can sleep; results wait in the bucket. You should see RUN_DIR=… and, at the end, the path to your new openjevx.w8.onnx under ~/openjevx/data/work/runs/.
The box fetches the training code from a pushed commit, so commit and push any change under finetuning/ before training. Your data goes through R2, not git.
6. Read the gate report
For each gate file you get three numbers:
- accuracy: the top answer was right.
- right & confident: right and sure enough to act on (yes ≥ 0.8, no ≤ 0.2, a choice or score ≥ 0.6). This is what matters: an unsure answer gets escalated, not acted on.
- confidently wrong: sure and wrong. The dangerous one; keep it near zero.
The model ships only if the everyday basics stay at least 90% right & confident, at most 2% confidently wrong, and it gets 12 of jevx's 13 fundamentals. These thresholds are in gate in ~/.config/openjevx/config.json. The full report is saved to ~/openjevx/data/work/gate/<model>.json. Compare your line with the baseline from step 4.
7. Run your model
Put the model in a folder with its settings and tokenizer:
my-model/
openjevx.w8.onnx
config.json (its own calibration temperatures)
tokenizer.json
Point the server at the folder in openjevx.json and start it:
{
"listen": "127.0.0.1:21118",
"device": "auto",
"model": "/path/to/my-model"
}
With Docker, mount the folder and set the same "model" path inside the container. Then add a jevx profile with a versioned model name, so cached answers from the old model are not reused:
jevx profile add mymodel http://127.0.0.1:21118/v1/systemone --model mymodel-v1
jevx profile use mymodel
jevx cache clear
jevx ask --noul discount="Does the order get the discount? Orders over Rs 500 get 10% off." --in '{"order_total_rs": 501}'
You should see a yes with high confidence. Train again later? Bump the name to mymodel-v2.
8. Troubleshooting
- The GPU box is stuck loading. Some hosts take long to start. After 15 minutes the box is destroyed and the next machine is tried, up to 3.
- Your network dropped. The box does not depend on your laptop; results wait in R2. Run the step again to pick them up.
- The trainer rejects a row. The CPU trainer check in step 5 names the row and source before any GPU is rented. Fix it in your CSV and import again.
- The gate fails. Look at which file failed. If it is your file, add near-miss rows and check the state holds every fact; if it is the basics, lower
--repeatso your data does not crowd them out.
Downloads
| What | File | Use it for |
|---|---|---|
| Fine-tune kit v0.4.0 (2.2 GB) | openjevx-finetune-v0.4.0.tar.gz (sha256) | scripts, trainable checkpoint (model.safetensors), openjevx.onnx, openjevx.w8.onnx, encoder/, tokenizer/; no training data |
| Trainable checkpoint v0.5.0 (777 MB) | openjevx-finetune-v0.5.0.tar.gz (sha256) | fine-tune further: model.safetensors, encoder/, tokenizer/, rl_agent_config.json; no training data |
| Model folder v0.5.0 (467 MB) | ask us for the download: [email protected] | run it: openjevx.w8.onnx + config.json (its calibration temperatures) + tokenizer.json; gate it (step 4) |
Pipeline (finetuning/ Taskfile, import_csv.py) | ask us for the download: [email protected] | steps 3 to 6 |
| Example CSV | decisions.csv | the starting point for step 2 |
The v0.5.0 run: one RTX 4090 on vast.ai, 734,145 decisions, one full pass in 2.4 h at about 88 items/s, about $1.10 in total. Test questions were removed first (22,592 leaked questions), and every row was checked against the trainer before renting.
Questions or a bug in the pipeline: write to [email protected].
Want this on your data?
Send us a sample of your decisions and we will run a pilot: [email protected].