Skip to main content
Liquid AI d1 model interface showing probability distribution, no output tokens. AI, machine learning, data science.

Editorial illustration for Liquid AI's d1 Model Returns Probabilities Without Output Tokens

Liquid AI's d1 Model Skips Text, Returns Pure Probabilities

Liquid AI's d1 Model Returns Probabilities Without Output Tokens

• 4 min read

Liquid AI shipped a model this week that never writes a word. Called d1, it's built to answer one narrow question: given this context and this fixed list of options, what's the probability of each one? No essays, no chain-of-thought, no generated tokens at all. The company put a number on it directly: usage.output_tokens reads 0 in every response.

That's a deliberate bet against how most teams currently handle classification work. Ticket routing, content moderation, reranking, LLM-as-judge scoring, these are jobs companies often hand to general-purpose LLMs, paying for token generation and prompt engineering to force a model into picking option A, B, or C. Liquid's pitch is that if the answer already lives inside a known set of outcomes, you don't need a language model composing sentences to get there. You need something that returns calibrated probabilities and stops.

d1 is live now on the Liquid API under the name d1:free, hosted only, no downloadable weights in GGUF, MLX, or ONNX format. Liquid's own migration guide frames the decision plainly: if the output is one of N known options, use a decision model. If it has to generate new text, stick with an LLM.

Liquid AI has released d1, a decision model built for structured choices instead of text generation. You give it context and a set of typed questions. It returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens.

Why this matters

For teams burning API calls on classification, moderation or routing, d1 is worth a look for one reason: it collapses work that used to take several LLM round trips into a single scored call. That 3-to-1 reduction in Liquid's own moderation example isn't a marketing flourish, it's the actual product. If those confidence thresholds hold up outside curated demos, that's real latency and cost saved for anyone running high-volume pipelines like ticket triage or content review.

We'd still want to see this stress-tested on messy, adversarial inputs before betting production infrastructure on it. Calibrated probabilities sound rigorous, but calibration claims from vendors deserve the same scrutiny as accuracy claims: ask for the eval data, not just the pitch. The fallback-to-bigger-model pattern when confidence dips below 0.5 is sensible engineering, but it also means d1 is explicitly not meant to replace your frontier model, just triage for it.

Watch whether independent benchmarks confirm the calibration holds across domains Liquid didn't test, and whether "zero output tokens" actually translates into the cost savings teams are hoping for once they factor in the hosted API pricing.

Common Questions Answered

How does Liquid AI's d1 model differ from traditional LLMs in handling classification tasks?

Unlike traditional LLMs that generate text output, d1 is specifically designed to return calibrated probabilities across a fixed set of outcomes without producing any generated tokens. The model takes context and a typed list of options as input, then directly returns probability scores for each option in a single call, eliminating the need for multiple LLM round trips.

What does it mean that d1 has zero output tokens in every response?

Zero output tokens means d1 never generates text or essays as part of its response. Instead of producing written content, it returns only structured probability values for the predefined options you provide. This is a fundamental architectural difference from generative models, making d1 more efficient for classification and decision-making tasks.

What are the practical use cases where d1 can reduce API costs and latency?

D1 is optimized for high-volume classification tasks like ticket routing, content moderation, LLM-as-judge scoring, and reranking. According to Liquid AI's moderation example, the model can collapse work that previously required three or more LLM calls into a single scored call, delivering significant latency and cost savings for teams running these types of pipelines.

Why is Liquid AI's approach with d1 considered a deliberate bet against current classification practices?

Most teams currently handle classification work by using general-purpose LLMs to generate chain-of-thought reasoning and essays, which requires multiple API calls and generates unnecessary tokens. D1 challenges this approach by proving that returning calibrated probabilities directly for a fixed set of options is more efficient, eliminating the overhead of text generation entirely.

LIVE03:02xAI's Recent Domain Transfer Fuels OpenAI 'Dots' Troll Speculation