Skip to main content
AI Agent Composer interface displayed on a screen, showcasing Contextual AI's top FACTS benchmark score. [allaccessible.org](

Editorial illustration for Contextual AI unveils Agent Composer, hits top FACTS benchmark score

Contextual AI Launches Agent Composer for Enterprise AI

Contextual AI unveils Agent Composer, hits top FACTS benchmark score

Updated: 3 min read

The new standard for AI trustworthiness just got a hard reset. Contextual AI didn’t just beat Google’s FACTS benchmark, they topped it, proving that grounded, hallucination-resistant outputs aren’t a pipe dream. The secret?

Fine-tuning Meta’s open-source Llama models on Vertex AI, targeting the very tendency of AI systems to invent. Now they’re packaging that rigor into Agent Composer, a platform that collapses complex engineering workflows into minutes. This isn’t another orchestration tool.

It’s a radical rethink: three distinct ways to build AI agents, all capable of coordinating multiple tools across multiple steps without the usual drift.

The approach has earned recognition. According to a Google Cloud case study, Contextual AI achieved the highest performance on Google's FACTS benchmark for grounded, hallucination-resistant results.

Agent Composer doesn’t just stitch together tools, it rewires the logic of enterprise AI deployment. By marrying Vertex AI’s infrastructure with fine-tuned Llama models that refuse to hallucinate, Contextual AI has solved the trust problem that has kept retrieval-augmented generation in the lab. The FACTS benchmark win is a signal, not a trophy: it proves that grounded, multi-step orchestration can scale without inventing answers.

For teams drowning in complex workflows, this is the difference between a proof of concept and a production system that actually holds up under audit. The race to build reliable agents just got a new finish line.

Common Questions Answered

What is the FACTS benchmark, and why is it important for Contextual AI's Grounded Language Model (GLM)?

The FACTS benchmark is a comprehensive evaluation tool designed to measure how accurately large language models ground their responses in provided source materials and avoid hallucinations. [deepmind.google](https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/) developed this benchmark with 1,719 carefully crafted examples to test long-form responses. Contextual AI's GLM achieved top performance on this benchmark, demonstrating its ability to minimize hallucinations and provide precise, attributable responses.

How does Contextual AI's Grounded Language Model (GLM) differ from traditional foundation models in handling enterprise AI applications?

Unlike traditional foundation models that may hallucinate or prefer their own parametric knowledge, the GLM is specifically engineered to minimize hallucinations in Retrieval-Augmented Generation (RAG) and agentic use cases. [contextual.ai](https://contextual.ai/blog/introducing-grounded-language-model) notes that the model provides inline attributions, directly citing the sources of retrieved knowledge within its responses. This approach addresses critical enterprise risks by delivering precise responses that are strongly grounded in specific retrieved source data.

What unique features does the FACTS Grounding dataset provide for evaluating language models?

The FACTS Grounding dataset includes 1,719 examples designed to test long-form responses, with documents ranging up to 32,000 tokens in length. [deepmind.google](https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/) created the dataset with a public set of 860 examples and a private set of 859 examples to prevent benchmark contamination. The dataset includes a comprehensive leaderboard to track progress in language model factuality and grounding.

LIVE13:01Scott Bessent Takes Aggressive Stance on Chinese AI