Skip to main content
IBM unveils Granite Speech 4.1 2B models, showcasing breakthrough 1.33 WER performance on LibriSpeech clean dataset, revoluti

Editorial illustration for IBM launches Granite Speech 4.1 2B models, hits 1.33 WER on LibriSpeech clean

IBM launches Granite Speech 4.1 2B models, hits 1.33 WER...

Updated: 3 min read

An autoregressive model that translates six languages. A non-autoregressive sibling that drops translation and Japanese entirely to shave latency. Both achieve a 1.33 Word Error Rate on LibriSpeech clean.

IBM’s Granite Speech 4.1 2B family is compact by design, 2 billion parameters, but the real story is in the fork. The standard model handles English, French, German, Spanish, Portuguese, and Japanese, plus bidirectional automatic speech translation. Its NAR counterpart strips out translation and Japanese, focusing exclusively on five languages for deployments where every millisecond matters.

That 1.33 WER is not the headline; it’s the baseline. The decision between autoregressive and non-autoregressive is the strategic one.

Granite Speech 4.1 2B is a compact and efficient speech-language model designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST) covering English, French, German, Spanish, Portuguese, and Japanese. Its non-autoregressive counterpart, Granite Speech 4.1 2B-NAR, focuses exclusively on ASR — specifically targeting latency-sensitive deployments — and supports English, French, German, Spanish, and Portuguese, but not Japanese. That’s a meaningful distinction: teams that need Japanese transcription or any speech translation capability should reach for the standard autoregressive model.

This isn’t just a benchmark number. It’s a tactical carve-out. IBM has drawn a clear line: the autoregressive model for Japanese transcription and translation; the NAR variant for everything else where milliseconds matter.

That 1.33 WER on LibriSpeech is impressive, but the real signal is the bifurcation. Teams chasing low-latency ASR in five languages now have a purpose-built tool, lean, fast, no translation overhead. For those needing the full multilingual stack plus Japanese, the heavier model is ready.

Bloat is the enemy of deployment. Granite Speech 4.1 sidesteps it with surgical precision. Good engineering chooses when to slow down and when to sprint.

IBM just gave developers a very clear stopwatch.

Common Questions Answered

What languages does the autoregressive Granite Speech 4.1 2B model support?

The autoregressive model supports six languages: English, French, German, Spanish, Portuguese, and Japanese. Additionally, it includes bidirectional automatic speech translation capabilities across these supported languages.

How does the non-autoregressive (NAR) variant of Granite Speech 4.1 2B differ from the standard model?

The NAR variant removes translation functionality and drops Japanese language support entirely to reduce latency for applications where milliseconds matter. This stripped-down version is optimized for low-latency automatic speech recognition in the remaining five languages.

What Word Error Rate does Granite Speech 4.1 2B achieve on LibriSpeech clean?

Both the autoregressive and non-autoregressive variants of Granite Speech 4.1 2B achieve a 1.33 Word Error Rate on LibriSpeech clean, which IBM positions as an impressive benchmark result for a compact 2 billion parameter model.

Why did IBM create two separate versions of the Granite Speech 4.1 2B model?

IBM created the bifurcated approach to serve different use cases: the autoregressive model handles scenarios requiring Japanese transcription and translation, while the NAR variant targets teams needing low-latency automatic speech recognition across five languages where processing speed is critical.

What is the parameter size of the Granite Speech 4.1 2B models?

The Granite Speech 4.1 2B family consists of models with 2 billion parameters, making them compact by design while still achieving competitive performance metrics on standard speech recognition benchmarks.

LIVE16:35Google Expands SynthID Watermark to Label AI Content