Skip to main content
Liquid AI unveils LFM2.5-230M model launch with open-source AI frameworks including llama.cpp, MLX, vLLM, SGLang, and ONNX fo

Editorial illustration for Liquid AI releases LFM2.5-230M, adds llama.cpp, MLX, vLLM, SGLang, ONNX

Liquid AI LFM2.5-230M Adds 5 Inference Engine Support

Liquid AI releases LFM2.5-230M, adds llama.cpp, MLX, vLLM, SGLang, ONNX

Updated: 3 min read

Liquid AI just released LFM2.5-230M. It is a small model, 230 million parameters, but it runs on almost anything: llama.cpp, MLX, vLLM, SGLang, ONNX. The benchmarks look like a misprint.

It beats Qwen3.5-0.8B and Gemma 3 1B IT on instruction following and data extraction. That is a 230M model outperforming models four times its size. The secret is distillation, the smaller model inherits behavior from the 350M sibling, then gets fine-tuned in three stages.

Honesty comes with the package: Liquid AI says don’t use it for reasoning, math, or code. But for large-scale data extraction, parsing 100,000 clinical reports into structured fields, this is exactly the right tool. And now it ships with support across the major inference engines, ready for on-device deployment.

Liquid AI shipped LFM2.5-230M , it’s the company’s smallest model to date. The release targets a specific job: running agentic tasks on phones, robots, and automation devices.

This is a model that knows exactly what it is, and what it isn’t. LFM2.5-230M doesn’t pretend to be a universal oracle. It’s a scalpel, not a sledgehammer.

On instruction following and structured data extraction, it punches far above its weight class, besting models several times its size. That’s the distillation advantage: targeted behavior inherited from a bigger sibling, without the bloat. And the framework support?

That’s the quiet revolution here. llama.cpp, MLX, vLLM, SGLang, ONNX, pick your stack. Deploy it on a phone, a laptop, or a low-power edge device.

The inference efficiency is built in, not bolted on. So don’t ask what this model can’t do. Ask what you actually need it to do.

If the answer is parse a million clinical reports, follow complex instructions, or extract structured fields from messy text, this is the tool. It’s honest about its limits. That honesty is the real strength.

In a field obsessed with bigger, faster, broader, Liquid AI built something useful precisely because it stayed small and focused. That’s not a compromise. That’s strategy.

Common Questions Answered

What specific model did Liquid AI release with the LFM2.5-230M designation?

Liquid AI released the LFM2.5-230M, a 230 million parameter language model that builds on their prior LFM2.5 architecture. This model is optimized for edge deployment and offers competitive performance despite its relatively small size.

Which inference frameworks are now supported by Liquid AI's LFM2.5-230M?

The LFM2.5-230M now supports llama.cpp, MLX, vLLM, SGLang, and ONNX frameworks. This broad compatibility allows developers to run the model on CPUs, Apple Silicon, and through high-performance serving systems.

What advantage does adding llama.cpp support provide for the LFM2.5-230M?

llama.cpp support enables efficient CPU inference, making the LFM2.5-230M accessible on a wide range of hardware without requiring a GPU. This integration is particularly useful for on-device applications and local deployments in resource-constrained environments.

How does vLLM integration benefit users deploying Liquid AI's LFM2.5-230M?

vLLM integration allows users to serve the LFM2.5-230M with high throughput and low latency using advanced batching and memory management techniques. This makes the model suitable for production-scale applications requiring real-time responses.

LIVE13:43Microsoft's New Coding AI Falls Short Against DeepSeek on Price and Performance