Deemwar ยท decision model DemoGamesRun it yourselfTraining bookRecipes

Is your AI bill
overkill?

Many calls to a big LLM are small decisions: which team owns this ticket, which action handles this request. Paste up to 50 of yours and see how many a small model you run yourself answers confidently, on a plain CPU, with your data staying yours.

Where to run it
Optional: estimate what this could save

Privacy: on our server, what you paste answers this request only; it is not logged or stored, and we keep only anonymous counts (visits, runs, clicks). Please don't paste secrets or personal data: it is a public demo. In-browser mode sends nothing anywhere. Limits: 50 lines, 3 checks per 10 minutes.

What we measured

Small numbers, stated with their method

598 MBthe 8-bit model file; about 1.6 GB of RAM once loaded; CPU only
178 msmedian for one yes/no (p95 192 ms) on an Apple M5 Pro laptop that was also busy. Our demo server is much slower.
36 / 40routing tickets right on our own labelled pool. 22/30 risky payments, 23/30 alerts. Small sets we wrote, recorded 1 Oct 2026.
"unsure"when it is not confident it says so instead of guessing. Those are the lines you keep on the big model.

We have no customer results yet and will publish only ones we can prove. See 18 recipes with real outputs, including where it misses.

How it works

Three steps, and you run the model

Bring your decisions

Policies, runbooks, tickets with the team that took them, alerts with whether anyone acted.

We train it, or you do

Labels come from rules we evaluate or outcomes that really happened, never from another model. You confirm every rule.

You run it

An 8-bit model on a plain CPU next to your production. You get a report on your own held-out cases: right and confident, confidently wrong, abstained.

Play

Can you beat the model?

Ten rounds against the clock, then compare with the model on the same ten. Shareable score card. Nothing you do here is stored.

Call it from code

One HTTP call

Against a server you run yourself (default port 21118). Real response, trimmed:

curl -s http://127.0.0.1:21118/v1/systemone -d '{
  "state": "I was charged twice for invoice 1042",
  "questions": {"q": {"type": "choice",
    "instructions": "Which team should handle this ticket?",
    "criteria": {"web": "frontend or UI", "api": "backend or API",
                 "billing": "payments and invoices", "docs": "how-to question"}}}}'

{"answers":{"q":{"choice":"billing","confidence":0.945,
  "probabilities":{"billing":0.945,"web":0.055}, ...}},"model":"openjevx"}

Question types: noul (yes/no), choice (pick one), score (rating). More in the recipes.

Get your own

Three ways, one question

Can your data leave your account?

Do it yourself, free

OpenJevX is open source (Apache-2.0) and Deemwar is its open-source partner. The training book walks from your decisions in a spreadsheet to an 8-bit model you run yourself; one fine-tune run on a rented GPU costs about $1 to $3 by the guide's own numbers. The public kit has the scripts and base model; a few pipeline files come from us on request ([email protected]).

Yes: we train it

We train on our GPUs from your documents and records, and hand you the model, a way to run it, and the report on your own cases. Retrain monthly. Status: pilot, run by hand for the first few teams.

No: train inside your cloud

A packaged appliance that trains in your own cloud account so nothing leaves it. Status: planned (DigitalOcean first, then AWS). Not available yet; ask for early access.

You always run the model

We charge for the hours we save you (data preparation, GPU work, checking), never for the code, and we never host your inference. This demo is marketing; it runs on our machine, your model runs on yours.

Get a pilot on your data

Send three decisions your team makes every day and where the data lives. We'll tell you honestly whether it fits, before anything is trained.

Email [email protected]