Skip to main content
AI monitoring beats testing for catching failures. Experts analyze data on screens, ensuring AI system reliability.

Editorial illustration for Monitoring Beats Testing for Catching AI Failures, Experts Say

Monitoring Beats Testing for AI Failures

4 min read

A single customer support conversation with an AI agent can read as flawless when scored on its own, then turn out to be a symptom of a product that's failing at scale. That contradiction is reshaping how companies test AI agents before shipping them. At VB Transform 2026, Harrison Chase of LangChain, Hui Zhang of Conviva, and Emmanuel Turlay of CoreWeave laid out a shift already underway in enterprise AI teams: away from grading individual transcripts and toward comparing groups of users against a baseline to catch failures that a single trace won't show.

Alongside that, teams are experimenting with smaller, cheaper judge models rather than defaulting to full-scale LLMs for every evaluation. Agent-as-judge, where one AI agent grades another's output, has gained traction but hasn't displaced LLM-as-judge, which Chase said is still the standard approach most teams reach for first. The bigger fault line, according to Zhang, isn't between those two automated methods.

It's between any form of automated grading and human review, a tradeoff Zhang said every company building on agents now has to confront directly.

A single AI agent conversation can look flawless scored on its own and still point to a broken product. That gap is driving a shift in how enterprises evaluate agents, away from scoring individual traces and toward comparing cohorts of users against a baseline.

Why this matters

For teams shipping agents right now, this is a real change in where the work goes. Pre-launch test suites still matter, but Chase's argument, backed by Zhang and Turlay's experience at Conviva and CoreWeave, is that they catch the failures you already anticipated. Cohort monitoring catches the ones you didn't. That's a harder sell internally, because it means budgeting engineering time for always-on checks in production rather than declaring victory after a clean eval run before launch.

We'd push back gently on one thing: this only works if someone owns the loop from online signal to offline test case. Skip that step and monitoring just becomes a dashboard nobody acts on. The upside, if teams do close that loop, is a system that gets sharper with every cohort of real users instead of decaying against a static benchmark. For founders under pressure to show eval scores to investors or customers, expect more scrutiny of what those scores actually predict about production behavior, and less patience for a green checkmark that hides a broken product.

Common Questions Answered

Why is monitoring AI agents more effective than traditional testing for catching failures?

Individual AI agent conversations can appear flawless when scored in isolation, but still indicate a broken product at scale. Monitoring catches failures that weren't anticipated during pre-launch testing, whereas traditional test suites only catch issues you already expected to find.

What is the shift in enterprise AI evaluation that Harrison Chase, Hui Zhang, and Emmanuel Turlay described at VB Transform 2026?

Enterprise teams are moving away from grading individual AI agent transcripts and toward comparing cohorts of users against a baseline. This shift prioritizes production monitoring over pre-launch test suites to identify unexpected failures that emerge at scale.

How does cohort monitoring differ from traditional pre-launch test suites for AI agents?

Pre-launch test suites catch anticipated failures based on known edge cases, while cohort monitoring compares groups of users against baselines to identify unexpected patterns and failures in production. Cohort monitoring requires ongoing engineering resources for always-on checks rather than declaring success after a clean evaluation run.

What internal challenge do teams face when implementing cohort monitoring for AI agents?

Teams must budget engineering time for continuous production monitoring rather than treating AI agent evaluation as complete after passing pre-launch tests. This represents a harder sell internally because it requires sustained resource allocation instead of a one-time validation process.

LIVE06:16Audit Framework Identifies Rater State Bias in RLHF Preference Data