Deemwar decision model · Recipes

Fine-tune it on your decisions

The shipped model does not know your rules: that your week is four days, that a discount starts above Rs 500, that repo billing belongs to the payments team. Write those decisions in a spreadsheet, and one command trains a model that does. Each step below says what you should see.

0. What you get and what it costs

You get your own openjevx.w8.onnx (8-bit, at most 750 MB) plus a trainable checkpoint. The model runs on your machine; your decisions never go to a hosted model.

Training runs on a rented GPU (an RTX 4090 on vast.ai, about $0.40–0.60 an hour). One run costs about $1–3. For reference, the v0.5.0 full run took about 2.4 hours on 734k decisions and cost about $1.10; the smoke run before it costs cents. The box deletes itself when it is done.

figurewhat it is
~$1–3one fine-tune run
~2.4 hv0.5.0 full run on a 4090
≤ 750 MByour 8-bit model

1. Install

You need Go, Task, uv and Python 3.11+ on your machine.

Get the public fine-tune kit: scripts, the trainable checkpoint and the ONNX model (v0.4.0, 2.2 GB). Verify the checksum before you unpack it:

curl -L -O https://pub-8da821f06ff747cda688f8267ed2aa96.r2.dev/openjevx-finetune-v0.4.0.tar.gz
curl -L -O https://pub-8da821f06ff747cda688f8267ed2aa96.r2.dev/openjevx-finetune-v0.4.0.tar.gz.sha256
shasum -a 256 -c openjevx-finetune-v0.4.0.tar.gz.sha256
tar -xzf openjevx-finetune-v0.4.0.tar.gz
cd openjevx-finetune-v0.4.0
go version && task --version && uv --version && python3 --version

You should see openjevx-finetune-v0.4.0.tar.gz: OK. The kit unpacks to model.safetensors (the trainable checkpoint), openjevx.onnx and openjevx.w8.onnx, encoder/, tokenizer/, and scripts/ (the trainer train_openjevx.py, quantize_w8.py, export_onnx_gpu.py, and the vast/ job scripts).

The one-command pipeline used in steps 3 to 6 (import_csv.py, the finetuning/ Taskfile with task cloud-setup, task gate and task all) is not inside the kit. Ask us for it: [email protected]. Commands below assume you have it in a folder named openjevx.

For the GPU step you also need:

Optional: the jevx client (for the 13-fundamentals check in the gate and for using the model afterwards; ask us for it) and Docker. Steps 2 to 4 need no GPU and no cloud account.

2. What goes in your CSV

A row is one decision you want the model to make.

Start from the example. It covers a 4-day week, 3 working hours a day, a late parcel, a discount over Rs 500, team routing by repo and ticket urgency: Download decisions.csv (39 rows).

The columns

columnrequired?what to putgood examplecommon mistake
stateyesthe facts, as a JSON object or plain text; include the threshold if it varies{"hours_logged_today": 3.1}just true, or leaving out the fact the rule needs
questionyesone decision, as you'd ask a colleagueIs the parcel late?two decisions in one: "is it late and who handles it"
typeno (noul)noul = yes/no, choice = pick one, score = ratingchoiceusing choice for yes/no
optionschoice, scorethe keys the model picks between, |-separated; optional short meaning after =; score levels lowest firstpayments=billing repos|web=frontend reposscore levels out of order (high|low|medium)
answeryesyes/no (also true/false, 1/0); one option key; a level name or index (0 = lowest)yes, payments, highan answer that is not one of the options
splitnotrain, test or gate; empty = 90% train, 10% gate, fixed per rowgateno gate rows, so you can't measure the result
sourcenoa tag for where the row came fromshop—

JSON in a CSV cell: wrap the cell in double quotes and double the quotes inside ("{""day"": ""Friday""}"). Any spreadsheet does this when you save as CSV.

One example per type: the CSV row, and what the model receives

Each row becomes one request to the server, POST /v1/systemone, with your state and question. Training teaches the model to give your answer to that request.

Yes/no (noul)

CSV row:

state,question,type,options,answer
"{""order_total_rs"": 501}","Does the order get the discount? Orders over Rs 500 get 10% off.",noul,,yes

Request the model receives:

{
  "state": {"order_total_rs": 501},
  "questions": {"q1": {"type": "noul",
    "instructions": "Does the order get the discount? Orders over Rs 500 get 10% off."}}
}

Pick one (choice)

CSV row:

state,question,type,options,answer
"{""repo"": ""billing"", ""title"": ""Invoice PDF missing tax line""}",Which team owns this?,choice,payments=billing and payments-core repos|platform=infra-terraform and api-gateway repos|web=web-frontend and admin-console repos,payments

Request the model receives:

{
  "state": {"repo": "billing", "title": "Invoice PDF missing tax line"},
  "questions": {"q1": {"type": "choice", "instructions": "Which team owns this?",
    "criteria": {"payments": "billing and payments-core repos",
                 "platform": "infra-terraform and api-gateway repos",
                 "web": "web-frontend and admin-console repos"}}}
}

Rating (score)

CSV row:

state,question,type,options,answer
"{""customers_affected"": 40, ""payments_failing"": false, ""workaround"": false}",How urgent is this ticket?,score,low|medium|high|critical,high

Request the model receives:

{
  "state": {"customers_affected": 40, "payments_failing": false, "workaround": false},
  "questions": {"q1": {"type": "score", "instructions": "How urgent is this ticket?",
    "criteria": ["low", "medium", "high", "critical"]}}
}

Where the facts come from

Put in the state exactly what your system will send the model when it asks for real.

sourcestate
a log alert{"line": "OOMKilled payments-api", "env": "prod", "errors_last_10m": 42}
a support ticket{"text": "Charged twice for order 8812", "customer_tier": "gold"}
a database record{"promised_by": "2026-10-03", "delivered_on": "2026-10-04"}
a pull request{"files_changed": 14, "touches": ["migrations/"], "summary": "adds a column"}
a form{"leave_days_requested": 6, "leave_days_left": 4}

Do and don't

dodon't
near-miss pairs at every threshold: 499 → no, 500 → no, 501 → yes; Thursday → yes, Friday → noput the answer inside the state ("eligible": true)
balance the answers: no single answer above 80% of a question's rowstrain on another model's guesses: labels come from the rule or from reality
20+ rows per questiongive the same state and question two different answers (a fact is missing)
keep a few of the hardest rows as split=gatecopy one wording everywhere; vary how the question and facts are phrased

How many rows

Start with 200–1,000 rows. Each rule becomes many rows: both sides of the threshold, different values, different wordings. The import step warns you when a question has fewer than 20 rows or one answer takes more than 80% of them. The example CSV is deliberately small (39 rows) to show the shape; it gets those warnings too.

3. Import and check

python3 finetuning/dataprep/import_csv.py yours.csv --name mydata --add-to-config

You should see the count of good and bad rows, then one line per question with its answers:

yours.csv: 39 good rows, 0 bad

question                                                     rows  labels
Does the order get the discount? Orders over Rs 500 get 10%     6  false=3, true=3
Which team owns this?                                           7  platform=3, payments=2, web=2
warning: 'Which team owns this?': only 7 rows; aim for 20+ with near-miss pairs around the rule
adapter: every row accepted
wrote    24 train rows -> ~/openjevx/data/train/mydata_train.jsonl
wrote    15 gate  rows -> ~/openjevx/data/gate/mydata_gate.jsonl
updated ~/.config/openjevx/config.json (backup: config.json.bak-…)

A bad row is reported with its line number and a fix (line 7: answer 'maybe' is not yes/no). Nothing is written until every row is good, unless you pass --skip-bad. Use --dry-run to check without writing.

--add-to-config backs up ~/.config/openjevx/config.json, then adds your train file to the training mix (repeated 3 times; change with --repeat), your gate file to the gate, and both to the leakage check. Re-running it does not add duplicates.

4. Test the current model on your gate first

Before paying for a GPU, see how the shipped model already does on your rules. You need the v0.5.0 model folder (openjevx.w8.onnx, config.json, tokenizer.json, 467 MB). It is not on a public download link yet: ask us for it at [email protected]. Unpack it next to the pipeline, then point the gate at it:

tar -xzf openjevx-model-0.5.0.tar.gz        # -> model/ (openjevx.w8.onnx, config.json, tokenizer.json)
cd finetuning
task gate -- ../model

This builds and serves the model on a free local port and scores every gate file, one line each:

gate/mydata_gate.jsonl   n=    15 accuracy  60.0% right&confident  40.0% confidently WRONG 13.3%

That line is your baseline. If your rules already score well, you may not need to train at all.

5. Train

cd finetuning
task all

It runs these stages and stops at the first failure:

  1. dataprep generates the built-in rule-labelled sets (basics, conditions, IT work, logs).
  2. leakage check removes any training question that also appears in a test or gate file.
  3. trainer check on CPU builds the shard and runs rows through the real trainer, so a data bug fails here, for free, not on a rented GPU.
  4. smoke run: about 10 minutes on a GPU to prove the whole path works.
  5. full run: the real training, a few hours.
  6. gate on the new model.

The GPU box downloads the shard from your private R2 bucket, trains, exports and quantizes to 8-bit, uploads the model, checkpoint and log back to R2, then destroys itself. A timer destroys it at the deadline no matter what. Your laptop only waits on R2, so it can sleep; results wait in the bucket. You should see RUN_DIR=… and, at the end, the path to your new openjevx.w8.onnx under ~/openjevx/data/work/runs/.

The box fetches the training code from a pushed commit, so commit and push any change under finetuning/ before training. Your data goes through R2, not git.

6. Read the gate report

For each gate file you get three numbers:

The model ships only if the everyday basics stay at least 90% right & confident, at most 2% confidently wrong, and it gets 12 of jevx's 13 fundamentals. These thresholds are in gate in ~/.config/openjevx/config.json. The full report is saved to ~/openjevx/data/work/gate/<model>.json. Compare your line with the baseline from step 4.

7. Run your model

Put the model in a folder with its settings and tokenizer:

my-model/
  openjevx.w8.onnx
  config.json        (its own calibration temperatures)
  tokenizer.json

Point the server at the folder in openjevx.json and start it:

{
  "listen": "127.0.0.1:21118",
  "device": "auto",
  "model": "/path/to/my-model"
}

With Docker, mount the folder and set the same "model" path inside the container. Then add a jevx profile with a versioned model name, so cached answers from the old model are not reused:

jevx profile add mymodel http://127.0.0.1:21118/v1/systemone --model mymodel-v1
jevx profile use mymodel
jevx cache clear
jevx ask --noul discount="Does the order get the discount? Orders over Rs 500 get 10% off." --in '{"order_total_rs": 501}'

You should see a yes with high confidence. Train again later? Bump the name to mymodel-v2.

8. Troubleshooting

Downloads

WhatFileUse it for
Fine-tune kit v0.4.0 (2.2 GB)openjevx-finetune-v0.4.0.tar.gz (sha256)scripts, trainable checkpoint (model.safetensors), openjevx.onnx, openjevx.w8.onnx, encoder/, tokenizer/; no training data
Trainable checkpoint v0.5.0 (777 MB)openjevx-finetune-v0.5.0.tar.gz (sha256)fine-tune further: model.safetensors, encoder/, tokenizer/, rl_agent_config.json; no training data
Model folder v0.5.0 (467 MB)ask us for the download: [email protected]run it: openjevx.w8.onnx + config.json (its calibration temperatures) + tokenizer.json; gate it (step 4)
Pipeline (finetuning/ Taskfile, import_csv.py)ask us for the download: [email protected]steps 3 to 6
Example CSVdecisions.csvthe starting point for step 2

The v0.5.0 run: one RTX 4090 on vast.ai, 734,145 decisions, one full pass in 2.4 h at about 88 items/s, about $1.10 in total. Test questions were removed first (22,592 leaked questions), and every row was checked against the trainer before renting.

Questions or a bug in the pipeline: write to [email protected].

Want this on your data?

Send us a sample of your decisions and we will run a pilot: [email protected].