Editorial illustration for Perplexity's 9B AI Model Scores 92.4% on MADQA Benchmark
Perplexity's 9B Model Hits 92.4% on MADQA Benchmark
Perplexity put two new embedding models on Hugging Face this week, and the size gap between them is the whole story. The smaller one, pplx-embed-v2-late at 0.6B parameters, is built to run on edge hardware and handle queries cheaply. The larger one, at 9B parameters, is the quality play, and Perplexity says it hits 92.4% on the MADQA benchmark. Both are ColBERT-style multimodal models, meaning they retrieve text, images, and rendered PDF pages out of the same embedding space, and both ship under an MIT license that permits commercial use.
The interesting bit is how the two sizes interact. Perplexity built them to share one embedding space, so a 9B index can be queried with the 0.6B model instead of its full-size counterpart. That matters for anyone trying to run retrieval at scale without paying 9B inference costs on every single query.
There's no hosted API yet, just Hugging Face checkpoints that need sentence-transformers 6.0.0 or newer and transformers 5.4.0 or newer to run, both confirmed for CUDA GPU use. A hosted endpoint is planned but not live.
Perplexity has released pplx-embed-v2-late, a pair of ColBERT-style multimodal embedding models. They come in 2 sizes: 0.6B for fast, cheap queries and 9B for maximum quality.
Why this matters
For teams building retrieval pipelines, the MIT license and same-day Hugging Face availability matter more than the headline 92.4% score. Perplexity isn't gatekeeping this behind an API you have to wait on; you can pull the weights today and run them against your own document store. The two-size split is the practical signal here: the 0.6B model, at roughly 340M active parameters, is clearly meant for production traffic where latency and cost dominate, while the 9B model is the one you reach for when you're benchmarking against BrowseComp+ or handling messy multimodal corpora like rendered PDFs and mixed image-text retrieval.
The 8.7-point margin on BrowseComp+ suggests dense embeddings are genuinely struggling with multi-hop or adversarial queries, which is worth testing against your own datasets rather than taking at face value. The weak spot, 61.2% on ViDoRe v3 Markdown, is a reminder that "retrieves text, images and PDFs" doesn't mean uniformly strong across formats. Anyone evaluating this for document-heavy RAG systems should benchmark on their own markdown-rendered content before assuming the top-line number generalizes.
Common Questions Answered
What are the two sizes of Perplexity's pplx-embed-v2-late models and what are they optimized for?
Perplexity released two versions: a 0.6B parameter model designed for edge hardware that prioritizes fast and cheap queries, and a 9B parameter model optimized for maximum quality that scores 92.4% on the MADQA benchmark. The smaller model is ideal for production environments where latency and cost are critical, while the larger model delivers superior performance for quality-focused applications.
What does ColBERT-style multimodal mean in the context of these embedding models?
ColBERT-style multimodal means that both pplx-embed-v2-late models can retrieve and embed multiple content types—text, images, and rendered PDF pages—from the same embedding space. This unified approach allows the models to handle diverse document formats and content types seamlessly within a single retrieval pipeline.
Why does Perplexity's MIT license and Hugging Face availability matter for teams building retrieval pipelines?
The MIT license and immediate Hugging Face availability mean that teams can access and deploy these models without waiting for API access or dealing with gatekeeping restrictions. Teams can pull the model weights directly and run them against their own document stores, giving them full control over their retrieval infrastructure and reducing dependency on external APIs.
What is the MADQA benchmark score achieved by Perplexity's 9B embedding model?
Perplexity's 9B parameter embedding model achieves a 92.4% score on the MADQA benchmark, demonstrating its high performance in retrieval and question-answering tasks. This benchmark score reflects the model's quality and effectiveness compared to other embedding models in the field.
Further Reading
- Multimodal embeddings beyond a single vector - Perplexity
- Perplexity's pplx-embed-v2-late Searches 190M Docs Without OCR or Chunking - AlphaSignal
- pplx-embed-v2-late-9b - Hugging Face
- pplx-embed-v2-late-0.6b - Hugging Face