Skip to main content
AI shopping agent on a laptop screen, biased product selection, ethical AI, consumer tech study.

Editorial illustration for Study: AI Shopping Agents Show 90-99% Bias in Product Selection

AI Shopping Agents Show 90-99% Bias in Picks

4 min read

Ask an AI shopping assistant to pick you a fitness watch, and the answer may hinge on nothing more than which article happened to load onto the screen first. That's the finding from a new study out of the Wharton School at the University of Pennsylvania, which put six current AI models through a test designed to mimic how agentic shopping tools actually work: reading a product page, pulling in outside sources, then making a call. The models included both mini and frontier-level versions, and the setup used something called the ACES simulator, short for Agentic e-Commerce Simulator, which feeds the AI a screenshot of a product grid and lets it reason its way to a pick.

The task itself was simple. Choose one fitness watch from a fixed set of options, acting as a personal shopping assistant would for a real customer. Researchers first checked what each model preferred with no outside input at all, then started introducing single sources, like a Reddit thread or a review site, to see how much that shifted the outcome. What they found suggests these agents are far more persuadable than a shopper might expect, and not evenly so across sources or models.

For shoppers, the study suggests that letting an AI agent buy on your behalf doesn't guarantee consistent or optimal purchase decisions. Two users with the same query, or the same user on a different day, can get different product recommendations with no visible reason.

Why this matters

For anyone building or deploying AI shopping agents, this study is a warning shot. A 90 to 99 percentage point swing in what a model recommends, triggered by nothing more than how a product listing is framed, isn't noise you can average away. It means the agent isn't reasoning about fitness watches at all.

It's reacting to whatever context happens to be sitting in front of it, and doing so with startling inconsistency between models like Claude Opus 4.8 and Gemini 3.5 Flash. The second experiment is the more unsettling result: stacking multiple sources should, in theory, dilute any single point of bias. Instead the distortion held.

That's a real problem for anyone pitching "autonomous purchasing" as a near-term product, and for retailers already worrying about how to get listed favorably in agent-driven searches. We'd treat any vendor claim about agent-based shopping reliability with real skepticism until someone shows these swings shrink under adversarial framing, not just clean lab grids. Worth watching whether Wharton or others test this against real e-commerce sites next, where the incentives to game an agent's context are already in place.

Common Questions Answered

What did the Wharton School study find about AI shopping agents and product selection bias?

The study found that AI shopping agents demonstrated 90-99% bias in product selection, meaning their recommendations were heavily influenced by factors like which product article loaded first rather than objective product quality. Six AI models were tested by simulating real agentic shopping workflows that involved reading product pages and pulling in outside sources to make recommendations.

Why do AI shopping agents produce inconsistent recommendations according to the research?

AI shopping agents produce inconsistent recommendations because they react to contextual framing rather than performing genuine reasoning about products. Two users with identical queries or the same user on different days can receive completely different product recommendations with no visible reason, indicating the agents are not making stable, logic-based decisions.

How does product listing framing affect AI model recommendations in shopping scenarios?

Product listing framing has a dramatic impact on AI recommendations, with a 90 to 99 percentage point swing in what models recommend based solely on how a product listing is presented. This inconsistency occurs across different frontier-level and mini AI models, including Claude Opus 4.8 and Gemini 3.5 Flash, suggesting the agents lack robust reasoning capabilities.

What are the practical implications of this AI shopping agent bias for consumers?

For consumers, the study indicates that relying on AI agents to make purchases on their behalf does not guarantee consistent or optimal purchase decisions. The lack of stable reasoning means shoppers cannot trust that they will receive the same recommendation twice or that the recommendation is based on genuine product analysis rather than arbitrary contextual factors.

LIVE22:07Barret Zoph Leaves OpenAI for Google After Five Months