Editorial illustration for New Jev Decision Index Ranks Open Reproductions at 64.1 Score
New Jev Decision Index Ranks Open Reproductions at 64.1...
The vLLM Semantic Router team put a number on something developers have been guessing at for months: how close open decision models actually get to the systems they're meant to replace. Decision 3.0 landed with five model sizes, from a 0.85B d3-lite up to a 26.09B flagship, all under Apache-2.0, all built to answer typed questions about text, JSON, and images without generating a single word of prose. Instead of a chatbot reply you parse for a verdict, these models hand back probabilities directly, Choice, Yes/No, or Score, scored in one pass.
That shift matters for anyone running a router or guardrail in production, where waiting on an LLM to write a sentence and then regexing it for a decision is slow and brittle. The lineup was tested on one AMD Instinct MI325X GPU in BF16, with no quantized builds listed yet and no disclosed context length. Scores on the internal board range from a near-perfect 97.7 on document extraction down to 34.1 on a harder vision benchmark, numbers the team is presenting as self-reported against live leaderboard data rather than independently verified.
Decision 3.0 multimodal decision models read text, JSON and images, then answer typed questions about them. There are 5 sizes, from 0.8B to 27B, all under Apache-2.0. For developers, this is a fast way to classify, route and gate requests without parsing generated text.
Why this matters
The Apache-2.0 license is the real headline here, not the 64.1 score. Five sizes spanning 0.85B to 26.09B means teams can actually test whether a 2.21B d3-nano handles their routing and gating tasks before committing GPU budget to the 26B flagship. That range matters more for production decisions than leaderboard position.
We'd treat the Jev Decision Index ranking itself with some caution. It's version 0.3.1, built by the same team releasing the model, and the card itself flags that Torchcast Decision 27B beats d3 on the public suite. That's a useful disclosure, but it also means the "ahead of Perplexity Decider v1.1 and Jev" framing deserves scrutiny rather than a straight repeat in your pitch deck.
For developers building classification or routing layers, the pitch is clear: skip parsing generated text, get typed answers directly from multimodal input. Whether that holds up outside benchmark conditions, on your own JSON schemas and image sets, is the thing to actually verify. Context length still isn't disclosed, which matters for anyone routing longer documents.
Further Reading
- Decision Index: reproduce the Jev decision-model benchmark suite - GitHub
- Decision 1.0 - vLLM Semantic Router
- Decision models explained: Jev vs Clef vs Strands Decider (2026) - eesel AI
- Introduce Decision 3.0 - vLLM Semantic Router