Skip to main content
Yandex Sona Recommender system diagram with dual attention mechanism for past and recent event processing.

Editorial illustration for Yandex Sona Recommender Uses Dual Attention for Past and Recent Events

Yandex Sona Recommender Uses Dual Attention for Past and...

• 3 min read

Yandex ran a seven-day live experiment on its smart speakers that replaced more than 15 candidate generators plus separate pre-ranking and ranking stages with a single transformer called Sona. The company's technical report frames this as a departure from how most production recommenders work: a cascade of models, each trained on its own objective, passing filtered candidates down the line until a heavy ranker makes the final call using hundreds of engineered features.

Yandex Music's previous stack leaned on exactly that setup, pulling in signals from Argus, an earlier recommender transformer, alongside hundreds of other hand-built features. Sona scraps the hand-engineering. It reads a listener's history once through a shared encoder, generates candidates with a decoder, and scores them with a Ranking Module that draws on the same encoder states. The only inputs are logged event fields like track ID, artist ID, duration, and played time, plus learned Semantic IDs.

The setting matters too: on smart speakers, playback often starts with no artist, genre, or mood selected by the user, what Yandex's researchers call a pure-recommendation problem. That's the backdrop for how Sona's architecture is built, starting with its semantic tokenizer.

Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex’s Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines.

Why this matters

Collapsing a multi-stage cascade into one generative model is the part worth watching, not the attention trick itself. Production recommenders at this scale usually run candidate generation, pre-ranking, and heavy ranking as separate systems with separate feature stores, which means separate teams and separate failure modes. Yandex is testing whether one model can do all three jobs on real traffic, on real devices, not just in an offline benchmark.

The 7-layer versus 1-layer split for recent versus older events is a pragmatic admission that full attention over a long history is too expensive to ship, so they're trading some quality for roughly half the inference cost. That's a tradeoff every team building sequential recommenders or long-context personalization has to make eventually, and seeing a named number attached to it is useful.

We'd want to see the actual A/B lift numbers and failure cases before calling this a cascade killer. A seven-day test on smart speakers is a narrow slice of traffic. Whether this generalizes to feeds, search, or ads ranking, where the engineered-feature cascade has decades of tuning behind it, is the real question.

LIVE10:30Yandex Sona Recommender Uses Dual Attention for Past and Recent Events