Skip to main content
Perplexity AI logo with a graph showing 8x smaller vectors, outperforming Voyage's embedding model.

Editorial illustration for Perplexity's New Embedding Model Outperforms Voyage With 8x Smaller Vectors

Perplexity's Embedding Model Beats Voyage 8x Smaller

• 4 min read

Perplexity Research and the vector database company turbopuffer have put out pplx-embed-v2-context-9b-preview, a contextual embedding model built for retrieval-augmented generation pipelines. It's live on Hugging Face under the MIT license, loadable with transformers 5.4.0 or later and trust_remote_code=True. It hasn't landed on the Perplexity API yet, and the model card warns that both the weights and the interface could change without keeping backward compatibility, so this is a preview in the literal sense.

The design answers a specific failure mode in how RAG systems chop documents into chunks. A chunk can depend on something stated pages earlier, a defined term, a heading, an entity introduced in a different section. Contextual embedding models have tried to fix this with late chunking, encoding the whole document first and pooling per chunk afterward.

But the training behind most of these models still picks a single "gold" chunk per query and treats everything else as a negative, even the supporting sentences a user would need to check the answer. Perplexity's team names three more cracks in that setup: binary labels are too coarse, LLM annotation costs scale badly, and labels get locked to one chunking scheme.

Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal.

Why this matters

For anyone running RAG in production, storage and latency are real line items, not abstractions. An 8x reduction in vector size, from 8 KB float32 down to 1 KB int8, while still edging out voyage-context-4 on retrieval quality, changes the math on what's affordable to index at scale. That matters more for founders watching infrastructure costs than for researchers chasing benchmark leaderboards, though the nDCG@10 numbers across 74 MTEB tasks give the latter something concrete to test against.

The training signal shift is worth watching closely. Retrieving an answer plus its supporting evidence, instead of a single gold passage, sounds like it should reduce the kind of confident-but-unverifiable outputs that make RAG systems hard to trust. Whether that holds up outside Perplexity's own suite is the open question.

MIT license and Hugging Face weights mean developers can actually kick the tires now, not wait for an API. The trust_remote_code requirement and transformers>=5.4.0 dependency are the kind of friction worth noting before you commit to a migration.

Common Questions Answered

How does pplx-embed-v2-context-9b-preview improve upon Voyage's embedding model?

Perplexity's new embedding model reduces vector size by 8x, from 8 KB float32 down to 1 KB int8, while still outperforming Voyage on retrieval quality metrics. This dramatic reduction in vector size significantly decreases storage and latency costs for production RAG systems without sacrificing performance.

What is the key innovation in how pplx-embed-v2-context-9b-preview processes documents for RAG pipelines?

The model uses contextual embeddings where each chunk is embedded with the full document in view, rather than in isolation. This contextual approach to the training signal represents the fundamental change that enables superior retrieval performance compared to traditional embedding methods.

Where can developers access pplx-embed-v2-context-9b-preview and what are the current limitations?

The model is currently available on Hugging Face under the MIT license and requires transformers 5.4.0 or later with trust_remote_code=True. However, it is still in preview status and has not yet been integrated into the Perplexity API, with the model card warning that both weights and interface may change without backward compatibility guarantees.

Why does the 8x reduction in vector size matter for production RAG deployments?

For companies running RAG systems at scale, storage and latency are significant operational costs rather than abstract concerns. The reduction from 8 KB to 1 KB per vector makes large-scale indexing substantially more affordable while maintaining or improving retrieval quality, directly impacting infrastructure budgets for founders and production teams.

How does pplx-embed-v2-context-9b-preview perform on standard benchmarks?

The model demonstrates competitive performance across 74 MTEB tasks, with nDCG@10 metrics showing it edges out Voyage-context-4 on retrieval quality. This benchmark performance provides concrete evidence of the model's effectiveness for researchers evaluating embedding models on leaderboards.

LIVE06:33Flow Engineering Raises USD 50M at USD 750M Valuation From Valor, Sequoia, Atreides