Editorial illustration for Contextual AI unveils Agent Composer, hits top FACTS benchmark score
Contextual AI Launches Agent Composer for Enterprise AI
Contextual AI unveils Agent Composer, hits top FACTS benchmark score
The new standard for AI trustworthiness just got a hard reset. Contextual AI didn’t just beat Google’s FACTS benchmark, they topped it, proving that grounded, hallucination-resistant outputs aren’t a pipe dream. The secret?
Fine-tuning Meta’s open-source Llama models on Vertex AI, targeting the very tendency of AI systems to invent. Now they’re packaging that rigor into Agent Composer, a platform that collapses complex engineering workflows into minutes. This isn’t another orchestration tool.
It’s a radical rethink: three distinct ways to build AI agents, all capable of coordinating multiple tools across multiple steps without the usual drift.
Agent Composer doesn’t just stitch together tools, it rewires the logic of enterprise AI deployment. By marrying Vertex AI’s infrastructure with fine-tuned Llama models that refuse to hallucinate, Contextual AI has solved the trust problem that has kept retrieval-augmented generation in the lab. The FACTS benchmark win is a signal, not a trophy: it proves that grounded, multi-step orchestration can scale without inventing answers.
For teams drowning in complex workflows, this is the difference between a proof of concept and a production system that actually holds up under audit. The race to build reliable agents just got a new finish line.
Common Questions Answered
What is the FACTS benchmark, and why is it important for Contextual AI's Grounded Language Model (GLM)?
The FACTS benchmark is a comprehensive evaluation tool designed to measure how accurately large language models ground their responses in provided source materials and avoid hallucinations. [deepmind.google](https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/) developed this benchmark with 1,719 carefully crafted examples to test long-form responses. Contextual AI's GLM achieved top performance on this benchmark, demonstrating its ability to minimize hallucinations and provide precise, attributable responses.
How does Contextual AI's Grounded Language Model (GLM) differ from traditional foundation models in handling enterprise AI applications?
Unlike traditional foundation models that may hallucinate or prefer their own parametric knowledge, the GLM is specifically engineered to minimize hallucinations in Retrieval-Augmented Generation (RAG) and agentic use cases. [contextual.ai](https://contextual.ai/blog/introducing-grounded-language-model) notes that the model provides inline attributions, directly citing the sources of retrieved knowledge within its responses. This approach addresses critical enterprise risks by delivering precise responses that are strongly grounded in specific retrieved source data.
What unique features does the FACTS Grounding dataset provide for evaluating language models?
The FACTS Grounding dataset includes 1,719 examples designed to test long-form responses, with documents ranging up to 32,000 tokens in length. [deepmind.google](https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/) created the dataset with a public set of 860 examples and a private set of 859 examples to prevent benchmark contamination. The dataset includes a comprehensive leaderboard to track progress in language model factuality and grounding.
Further Reading
- Contextual AI Announces General Availability of Its Enterprise Platform — Contextual AI Blog
- Contextual AI Launches World's First Instruction-Following Reranker — Contextual AI Blog
- 2026 data predictions: Scaling AI agents via contextual intelligence — SiliconANGLE