RSI-Jev v5.0-VL 3B
Fast, intuitive decisions: one pass, no text generated, a calibrated probability for every answer.

Which speed limit number is written on the object in the image?
System 2: writes, then you parse.{"speed_limit": "80"}chat model · 11 tokens · 333 ms · no probability
Recorded with v4.0-VL on one GB10. Replayed in the race below.
POST /v1/systemone
A language model is built to write; its deepest layers mostly shape the next word. A System One model never writes. It reads, then decides. So v5.0-VL keeps the first 20 of Qwen3.5-4B-Base’s 32 layers: 3.25B parameters, one pass.
Trained once per exit layer with the same recipe, the cut at 20 scores like all 32 on the decision suite (0.760 vs 0.761) and the held-out set (0.688 vs 0.686). Knowledge-heavy MMLU-Pro is the exception: 0.422 vs 0.457.
Best at 3B and under on the public board (2026-09-28), and ahead of eleven 4B-class entries. Sizes are served parameters. Our v5.0-VL score is a full run of the suite served by rsi-jev serve; v4.0-VL’s is its submitted run.
It picks the answer, with a confidence. Send one to four images with a request: a photo, a screenshot, a chart. Ask the same typed questions as with text.
36 decisions · median 76 ms each
Every answer here is a recorded output, misses marked ✗. Across all 143 items we ran (89 images), 187 of 227 answers were right.
Hover or tap a picture for the full question and every option. Times are whole calls on one GB10. Images: VisA (Zou et al. 2022) and CLEVR (Johnson et al. 2017), CC BY 4.0, resized; road and object photos from Wikimedia Commons via reinhard-z/vision-jev: stop sign, speed sign and bicycle public domain (Dori, P.Ctnt, Ingolfson), teddy, dog, car, bag and pedestrians CC0, boy with ball CC BY 2.0 (Evgeniy Isaev), red and green lights and leaves CC BY-SA 3.0 (Jacklee; Wald1siedel), cat CC BY-SA 4.0 (Grendelkhan). The crops of the CC BY-SA photos are shared under the same licences. Charts, screens, receipt and shapes are ours (CC0). Breakout from hr98w/jev-visual (MIT). Full credits.
Every frame it gets one screenshot and says which of five lanes the ball is in, and the paddle moves there.
Recorded live, real speed: 79 ms per round trip, 200 decisions. Lane probabilities from the game’s own feed, cropped from the same recording. Game from hr98w/jev-visual (MIT), unchanged.
rsi-jev demo opens this page on your machine.
Recorded live: child in the photo, yes 99.8%; the place, beach 99.8%; closest to the camera, the ball 89%. 90 ms on the server, one GB10. Photo: child.jpg by Evgeniy Isaev, CC BY 2.0.
RSI-Jev and a chat model.

RSI-Jev’s time is its median over 9 calls; the chat model’s text appears at the token times of one streamed call (greedy, up to 256 tokens). Which one is more accurate is unresolved: most of the chat model’s invalid replies gave an answer under the wrong JSON key; read leniently, it gets 23 of 30 right to our 26, a gap 30 questions cannot settle. Leaves photo by Wald1siedel (CC BY-SA 3.0; the resized copy is shared under CC BY-SA 3.0). The receipt is generated; its blank lower half is cropped here.








One leaking capsule in the tray, and it has never seen a defect example
Production-line inspection photo of green gel capsules.
Is everything in the photo good, or is at least one part defective?
✓ matches the reference254 ms on one GB10
A burnt spot on a circuit board
Production-line inspection photo of a printed circuit board.
Is everything in the photo good, or is at least one part defective?
✓ matches the reference245 ms on one GB10
Red light ahead: wait for green
What should the car do about the object in the image, which is on the road, in the car's lane, ahead of the car?
✓ matches the reference65 ms on one GB10
Reads the sign and changes speed
What should the car do about the object in the image, which is on the road, in the car's lane, ahead of the car?
✓ matches the reference65 ms on one GB10
A teddy bear on a bench, and no child
Is there a teddy bear in the image?
✓ matches the reference60 ms on one GB10
Which plan ends cheaper?
Which plan is cheaper in the last month shown?
✓ matches the reference117 ms on one GB10
The user said keep the file
The user said: "Actually, keep that file."
Which button should the agent click?
✓ matches the reference79 ms on one GB10
Card declined: what next?
Goal: finish paying for the order.
What happened to the payment?
✓ matches the reference121 ms on one GB10
Real output from v4.0-VL on one GB10; the time is the whole call. The three Wikimedia photos are shown at higher resolution than the 512 px copies the model was sent. Photos: VisA (Zou et al. 2022, CC BY 4.0, resized); Wikimedia Commons speed sign by P.Ctnt (public domain) and traffic light by Jacklee (CC BY-SA 3.0; the resized red-light photo is shared under CC BY-SA 3.0) and teddy bear via Pixabay (CC0). The charts and screens are generated by us.
AUROC. Bars start at 50, the score of a coin flip.
VisA: the same 2,162 photos of 12 kinds of parts for all three, one photo at a time, no defect examples. Score is AUROC: 50 is a coin flip, 100 is perfect. Ours 95% CI 85.1–88.1. The other two are scored from their authors’ published predictions.
RSI-Jev gives each part a probability that it is defective. The factory decides where to cut. Move the line and see what it catches.
Each dot is one photo’s served probability from the benchmark run behind the AUROC above, one call per photo. At 50% the line catches 405 defects and wrongly flags 34 good parts. The belt shows 12 of those photos with their real scores. Photos: VisA, Zou et al. 2022, CC BY 4.0, resized.
Same API as TypeSafe’s Jev, so existing Jev code works. Not affiliated with TypeSafe.
The chart runs to v3.0. v4.0-VL then added two rounds of image training and kept the text score at 0.756. Its last stage is RL that penalises confident mistakes: held-out calibration error before any post-hoc fix fell from 0.200 to 0.082 (one seed). v5.0-VL cuts a 4B model at layer 20 and retrains its decision head to read each option from its own tokens: Decision Index 38.38.
The experiments were run by the next generation of AutoScientists, our AI research-agent system. Read the experiment log: every arm through v3.0, and every arm on the v4.0-VL and v5.0-VL lines.
“My card was charged twice for order #4521. Please refund the second charge, I need the money back before rent is due on Friday.”
Real output, in about 35 ms on a desk-side GPU. It never writes free text, so there is nothing to parse and nothing made up. (Measured on v3.0.)
pip install "rsi-jev[fast,vision] @ git+https://github.com/Shanghua-Gao/RSI-Jev"rsi-jev serve v5.0-vl-3bcurl localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"state": [{"role": "user", "content": "I was charged twice. Please refund."}],
"questions": {"refund": {"type": "noul",
"instructions": "Does the user request a refund?"}}}'Already using Jev? Point your client at http://localhost:8000 instead. For an image, add an images list; the README shows how.
Recorded output, measured on v3.0.
The support ticket above: one request, three questions. Median of several runs (measured on v3.0).
Jev-style models answer in one step instead of writing text, which is why they are all this fast. Jev’s time includes the trip over the internet from our machine. The chat model is Qwen3.5-2B, the same size as RSI-Jev, writing its answer as JSON.
We kept working on how it runs. The same model on the same machine now answers 1.5–5.8× faster, with the same answers.
So you can use the confident answers directly and send the unsure ones to a person. Measured on v3.0, on 2,000 real test questions.
Built and measured on an HP ZGX Nano. Thanks to HP and NVIDIA for providing the HP ZGX Nano AI Station, powered by the NVIDIA GB10 Grace Blackwell Superchip.