RSI-Jev

RSI-Jev v5.0-VL 3B

A System One model.

Fast, intuitive decisions: one pass, no text generated, a calibrated probability for every answer.

A round road sign showing a speed limit of 80.
image + question

Which speed limit number is written on the object in the image?

probabilities61 ms
8099.7%
50<1%
none<1%
2 more answers<1%

System 2: writes, then you parse.{"speed_limit": "80"}chat model · 11 tokens · 333 ms · no probability

Recorded with v4.0-VL on one GB10. Replayed in the race below.

POST /v1/systemone

System One doesn’t need the whole LLM.

A language model is built to write; its deepest layers mostly shape the next word. A System One model never writes. It reads, then decides. So v5.0-VL keeps the first 20 of Qwen3.5-4B-Base’s 32 layers: 3.25B parameters, one pass.

Trained once per exit layer with the same recipe, the cut at 20 scores like all 32 on the decision suite (0.760 vs 0.761) and the held-out set (0.688 vs 0.686). Knowledge-heavy MMLU-Pro is the exception: 0.422 vs 0.457.

Decision Index 0.2.1

RSI-Jev v5.0-VL, 3B38.38
Intern-Decision-4B, 4.7B37.81
NeoHorse-Jev-4B, 4.7B36.75
Kev 4B, 4.7B34.64
Decider 2B, 2.3B28.97
RSI-Jev v4.0-VL, 2.3B28.31

Best at 3B and under on the public board (2026-09-28), and ahead of eleven 4B-class entries. Sizes are served parameters. Our v5.0-VL score is a full run of the suite served by rsi-jev serve; v4.0-VL’s is its submitted run.

Show it a picture.

It picks the answer, with a confidence. Send one to four images with a request: a photo, a screenshot, a chart. Ask the same typed questions as with text.

36 decisions · median 76 ms each

Every answer here is a recorded output, misses marked ✗. Across all 143 items we ran (89 images), 187 of 227 answers were right.

Hover or tap a picture for the full question and every option. Times are whole calls on one GB10. Images: VisA (Zou et al. 2022) and CLEVR (Johnson et al. 2017), CC BY 4.0, resized; road and object photos from Wikimedia Commons via reinhard-z/vision-jev: stop sign, speed sign and bicycle public domain (Dori, P.Ctnt, Ingolfson), teddy, dog, car, bag and pedestrians CC0, boy with ball CC BY 2.0 (Evgeniy Isaev), red and green lights and leaves CC BY-SA 3.0 (Jacklee; Wald1siedel), cat CC BY-SA 4.0 (Grendelkhan). The crops of the CC BY-SA photos are shared under the same licences. Charts, screens, receipt and shapes are ours (CC0). Breakout from hr98w/jev-visual (MIT). Full credits.

It plays Breakout from pixels.

Every frame it gets one screenshot and says which of five lanes the ball is in, and the paddle moves there.

Recorded live, real speed: 79 ms per round trip, 200 decisions. Lane probabilities from the game’s own feed, cropped from the same recording. Game from hr98w/jev-visual (MIT), unchanged.

Drop in your own photo and ask.

rsi-jev demo opens this page on your machine.

Recorded live: child in the photo, yes 99.8%; the place, beach 99.8%; closest to the camera, the ball 89%. 90 ms on the server, one GB10. Photo: child.jpg by Evgeniy Isaev, CC BY 2.0.

Same picture, same question.

RSI-Jev and a chat model.

Replay of measured timings on one GB10
A round road sign showing a speed limit of 80.

RSI-Jev v4.0-VL2B, one pass0 ms
Qwen3.5-2B chat2B, writes its answer as JSON0 ms
All 30 questionsRSI-JevChat model Median time69 ms540 ms Valid answers30 / 3020 / 30 Probability per answeryesno, text only

RSI-Jev’s time is its median over 9 calls; the chat model’s text appears at the token times of one streamed call (greedy, up to 256 tokens). Which one is more accurate is unresolved: most of the chat model’s invalid replies gave an answer under the wrong JSON key; read leniently, it gets 23 of 30 right to our 26, a gap 30 questions cannot settle. Leaves photo by Wald1siedel (CC BY-SA 3.0; the resized copy is shared under CC BY-SA 3.0). The receipt is generated; its blank lower half is cropped here.

A probability for every answer.

A tray of green gel capsules; one of them is leaking.
VisA, Zou et al. 2022, CC BY 4.0, resized

One leaking capsule in the tray, and it has never seen a defect example

Production-line inspection photo of green gel capsules.

Is everything in the photo good, or is at least one part defective?

68%Defective
  • Defective68%
  • All good32%

✓ matches the reference254 ms on one GB10

Real output from v4.0-VL on one GB10; the time is the whole call. The three Wikimedia photos are shown at higher resolution than the 512 px copies the model was sent. Photos: VisA (Zou et al. 2022, CC BY 4.0, resized); Wikimedia Commons speed sign by P.Ctnt (public domain) and traffic light by Jacklee (CC BY-SA 3.0; the resized red-light photo is shared under CC BY-SA 3.0) and teddy bear via Pixabay (CC0). The charts and screens are generated by us.

Finding defects on factory parts, with no defect examples.

RSI-Jev v4.0-VL, 2B86.6
Gemma 4 12B82.9
Jev-Omni (Gemma 4 12B + head)81.1

AUROC. Bars start at 50, the score of a coin flip.

VisA: the same 2,162 photos of 12 kinds of parts for all three, one photo at a time, no defect examples. Score is AUROC: 50 is a coin flip, 100 is perfect. Ours 95% CI 85.1–88.1. The other two are scored from their authors’ published predictions.

Where would you set the line?

All 2,162 VisA test photos, recorded scores

RSI-Jev gives each part a probability that it is defective. The factory decides where to cut. Move the line and see what it catches.

50%
405/ 1,200defective parts caught
34/ 962good parts wrongly flagged
defective part (1,200)good part (962)bright dots are flagged at this line
RSI-JEV LOOKS HERE

Each dot is one photo’s served probability from the benchmark run behind the AUROC above, one call per photo. At 50% the line catches 405 defects and wrongly flags 34 good parts. The belt shows 12 of those photos with their real scores. Photos: VisA, Zou et al. 2022, CC BY 4.0, resized.

6releases in 8 days, v1.0 to v5.0-VL
300+experiments by the agents, failures included
13 / 14same answers as Jev in our side-by-side test (measured on v3.0)

Same API as TypeSafe’s Jev, so existing Jev code works. Not affiliated with TypeSafe.

How the agents build it

1. ProposeAn agent suggests a change: new training data, a new loss, a new architecture.
2. Predict, then testIt writes down what it expects before spending any GPU time, then runs the experiment.
3. Keep only what winsA change ships only if it beats the current best on the test suite. Everything else is published as a failure.
The champion line climbs from v1.0 (0.622) to v2.0 (0.709), v2.1 (0.736) and v3.0 (0.756) on the 15-benchmark suite; grey dots are the experiments tried in between; below, the propose, experiment, learn loop. The champion line climbs from v1.0 (0.622) to v2.0 (0.709), v2.1 (0.736) and v3.0 (0.756) on the 15-benchmark suite; grey dots are the experiments tried in between; below, the propose, experiment, learn loop.

The chart runs to v3.0. v4.0-VL then added two rounds of image training and kept the text score at 0.756. Its last stage is RL that penalises confident mistakes: held-out calibration error before any post-hoc fix fell from 0.200 to 0.082 (one seed). v5.0-VL cuts a 4B model at layer 20 and retrains its decision head to read each option from its own tokens: Decision Index 38.38.

The experiments were run by the next generation of AutoScientists, our AI research-agent system. Read the experiment log: every arm through v3.0, and every arm on the v4.0-VL and v5.0-VL lines.

You ask. It picks the answer, with a confidence.

Customer message

“My card was charged twice for order #4521. Please refund the second charge, I need the money back before rent is due on Friday.”

RSI-Jev’s answers

Real output, in about 35 ms on a desk-side GPU. It never writes free text, so there is nothing to parse and nothing made up. (Measured on v3.0.)

Three commands

1
pip install "rsi-jev[fast,vision] @ git+https://github.com/Shanghua-Gao/RSI-Jev"
2
rsi-jev serve v5.0-vl-3b
3
curl localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": [{"role": "user", "content": "I was charged twice. Please refund."}],
  "questions": {"refund": {"type": "noul",
    "instructions": "Does the user request a refund?"}}}'

Already using Jev? Point your client at http://localhost:8000 instead. For an image, add an images list; the README shows how.

More examples

Recorded output, measured on v3.0.

How long does an answer take?

The support ticket above: one request, three questions. Median of several runs (measured on v3.0).

Jev-style models answer in one step instead of writing text, which is why they are all this fast. Jev’s time includes the trip over the internet from our machine. The chat model is Qwen3.5-2B, the same size as RSI-Jev, writing its answer as JSON.

And faster than when v3.0 first shipped

We kept working on how it runs. The same model on the same machine now answers 1.5–5.8× faster, with the same answers.

Milliseconds per request for six workloads, v3.0 as first released against now: 1.5x to 5.8x faster. Milliseconds per request for six workloads, v3.0 as first released against now: 1.5x to 5.8x faster.

How we made it faster, and what didn’t work

When it says it’s 90% sure, it’s right 99.5% of the time.

So you can use the confident answers directly and send the unsure ones to a person. Measured on v3.0, on 2,000 real test questions.

Built and measured on an HP ZGX Nano. Thanks to HP and NVIDIA for providing the HP ZGX Nano AI Station, powered by the NVIDIA GB10 Grace Blackwell Superchip.