Editorial illustration for Expedia AI chief: Users must have final say over AI agents
Expedia AI Chief: Users Must Control AI Agents
Xavi Amatriain has a new definition for the product requirements document. At VB Transform 2026 in Menlo Park last week, Expedia Group's first chief AI and data officer told the audience that evals, the tests used to grade AI system performance, have effectively replaced the PRD as the place where product intent gets written down. Red-teaming, security requirements, the whole spec: it goes into the evals now, before a line of code exists.
Amatriain took the job in December 2025 after running AI and Compute Enablement at Google, where he worked across the infrastructure behind Gemini and Search. He's also mentored people who went on to start Perplexity and Scale AI, so his read on where AI development is headed carries some weight.
The stakes behind that shift are real. VentureBeat's VB Pulse research surveyed 157 enterprises and found 66% either already let AI agents into production without human review or plan to within a year. Only 5% said they fully trust the automated evals making that call. Half have shipped an agent that cleared internal testing, then failed in front of an actual customer.
“The new PRD are the evals,” Xavi Amatriain, Expedia Group’s first chief AI and data officer, told the VB Transform 2026 audience last week in Menlo Park.
Why this matters
Amatriain's framing puts a stake in the ground at a moment when a lot of teams are still treating "agentic" as a checkbox feature rather than a design decision with consequences. Evals-as-PRD is a useful reframe for builders: if you're not writing down what "good" looks like, including adversarial cases, before you write code, you're shipping intentions instead of specifications. That's a discipline problem as much as a technical one.
The click-to-confirm rule matters more for what it signals than for the UX pattern itself. Expedia is a company that lets an AI agent book flights and hotels, real money, real itineraries, and its own AI chief is drawing a hard line at final user consent. For founders racing to ship autonomous agents, that's a useful check: agency isn't a nice-to-have you bolt on later, it's a security boundary that shapes what your evals even need to test. Worth watching whether this becomes an industry norm or stays an Expedia-specific caution.
Common Questions Answered
What does Xavi Amatriain mean by 'evals are the new PRD' at Expedia?
Amatriain, Expedia Group's first chief AI and data officer, argues that evals (tests used to grade AI system performance) have replaced traditional product requirements documents as the place where product intent is documented. Evals now include red-teaming, security requirements, and the complete specification before any code is written, fundamentally changing how AI products are defined and built.
Why does treating AI agents as a checkbox feature create consequences according to the article?
The article emphasizes that agentic AI should be treated as a design decision with real consequences, not just a feature to check off. When teams don't write down what 'good' looks like—including adversarial cases—before coding, they end up shipping intentions instead of specifications, which represents a discipline problem as much as a technical one.
How does the evals-as-PRD approach change the product development process for AI systems?
Instead of writing traditional product requirements documents first and then building code, the evals-as-PRD approach requires teams to define specifications, security requirements, and adversarial test cases upfront through evaluations. This ensures that product intent is clearly specified before development begins, creating a more disciplined and intentional approach to building AI products.
When did Xavi Amatriain take his role as Expedia Group's chief AI and data officer?
Xavi Amatriain took the position of Expedia Group's first chief AI and data officer in December 2025. He shared his insights about the importance of evals-as-PRD at the VB Transform 2026 conference in Menlo Park.
Further Reading
- Papers with Code - Latest NLP Research - Papers with Code
- Hugging Face Daily Papers - Hugging Face
- ArXiv CS.CL (Computation and Language) - ArXiv