Editorial illustration for Liquid AI's New d1-3B Model Processes Images in 35 ms for Edge Inspection
Liquid AI's d1-3B Processes Images in 35ms for Edge
Liquid AI put a number on its new vision model Tuesday: 35 milliseconds to process an image on the NVIDIA Jetson stack, fast enough for a factory camera checking parts on a line without ever generating a sentence. The company's latest release, Open d1, swaps out the usual chatbot playbook entirely. Instead of writing text token by token, the two models in this family, d1-3B and d1-omni-600M, take in a state and a set of named questions and spit back a probability distribution in a single forward pass. Every response logs output_tokens: 0.
d1-3B, built on LFM2.5-VL-3B with 3.12 billion parameters, handles text and images. d1-omni-600M, an early research checkpoint with no published latency numbers yet, handles text paired with either an image or audio. Both are open-weight, sitting on Hugging Face with Transformers support and day-one compatibility with llama.cpp, and both ship under the LFM Open License v1.0, which permits free commercial use for companies under $10 million in annual revenue.
Liquid AI is positioning these as tools for routing, moderation, and guardrail checks, not conversation. The architecture behind that claim is where things get specific.
A generative LLM writes its answer token by token, and your code parses it. A decision model takes a state and a set of named questions. It reads them once and returns a probability for every allowed answer.
Why this matters
For teams building on Jetson or RTX hardware, the pitch here is narrow and specific: skip the token generation step entirely when the task is really a classification problem wearing a chatbot's clothes. A 35 ms response on Jetson AGX Thor for a 384px frame is fast enough for frame-by-frame camera work, which is why Liquid AI is aiming d1-3B at visual inspection and moderation rather than general chat. That's a sensible bet.
Most "AI" deployed on a factory floor or a retail camera doesn't need prose, it needs a calibrated yes/no/this-category answer, and paying the latency and compute cost of an LLM for that is wasteful. The open-weight release on Hugging Face means developers can actually test the 35 ms claim themselves rather than taking Liquid AI's word for it, which matters given how often edge-AI benchmarks turn out to be best-case numbers. The real test is whether d1-omni-600M's voice routing and d1-3B's gesture control hold up outside curated demos, on noisy camera feeds and real microphones.
Worth watching before anyone puts this in production.
Common Questions Answered
How fast does Liquid AI's d1-3B model process images on NVIDIA Jetson hardware?
Liquid AI's d1-3B model processes images in just 35 milliseconds on the NVIDIA Jetson stack, specifically achieving a 35 ms response time on Jetson AGX Thor for 384px frames. This speed is fast enough for real-time frame-by-frame camera work in factory inspection applications without requiring token generation.
What is the key difference between Liquid AI's decision models and traditional generative LLMs?
Unlike generative LLMs that write answers token by token, Liquid AI's d1 decision models take a state and a set of named questions, then return a probability distribution for every allowed answer in a single forward pass. This approach eliminates the token generation step entirely, making it more efficient for classification tasks.
What are the two models in Liquid AI's d1 family and what are their specifications?
Liquid AI's d1 family consists of two multimodal decision models: d1-3B and d1-omni-600M. These models are designed specifically for visual inspection and moderation tasks rather than general chat applications, making them ideal for deployment on factory floors and retail environments.
Why is Liquid AI targeting the d1-3B model at visual inspection rather than general chat applications?
Liquid AI is positioning d1-3B for visual inspection and moderation because the 35 ms processing speed on Jetson hardware is fast enough for real-time frame-by-frame camera work on factory floors and retail environments. The decision model approach is specifically optimized for classification problems, making it more practical for these use cases than general-purpose chatbots.
Further Reading
- Liquid AI releases open d1 models for multimodal edge decisions - RuntimeWire
- Liquid AI d1 Vision: 19-200x Cheaper Than GPT-6.1 Sol - Explainx.ai
- Introducing d1: The most capable decision model, now with vision - Liquid AI
- LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge - Liquid AI
- Vision Models - Liquid AI Documentation