Skip to main content
Onton AI search platform showing 2.7x more accurate results than e-commerce engines, enhancing online shopping.

Editorial illustration for Onton Claims 2.7x More Accurate Search Than Top E-commerce Engines

Onton's AI Search 2.7x More Accurate Than E-commerce Giants

4 min read

Onton put a number on something most shoppers already know from experience: site search on the big platforms often misses what you actually mean. The San Francisco company's new model, Ontology 1, scored a mean precision@10 of 0.630 on a 90-query benchmark judged by three independent LLMs, against 0.543 for Google Shopping and 0.469 for Amazon. It managed that while indexing only about 1% of the catalog size those platforms work with, a gap that points to a different approach rather than brute-force scale.

Ontology 1 is neurosymbolic, built for conversational and multimodal queries that don't reduce cleanly to category and attribute filters. It's live now at Onton.com, with partner access granted case by case rather than through an open API or downloadable checkpoint. Onton is positioning this for agentic commerce and for retailers whose existing search stacks lose ground on long, requirements-heavy queries. Right now it only indexes home decor and furniture, though Onton says the underlying method extends to other verticals and to non-product data without much reconfiguration.

The gap shows up clearest on the kind of query that has no clean filter attached to it.

Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query benchmark scored by three independent LLM judges, Ontology 1 reached a mean precision@10 of 0.630, against 0.543 for Google Shopping and 0.469 for Amazon. It did this while indexing roughly 1% of their catalogs.

Why this matters

The gap between 0.630 and 0.543 sounds decisive until you remember Ontology 1 indexed about 1% of the catalog that Amazon and Google Shopping are working against. That's not a fair fight, it's a smaller pond. Precision@10 gets easier when you've pre-curated the water.

We'd want to see the benchmark rerun at 10% or 50% catalog coverage before treating this as proof of a better retrieval architecture rather than a better-groomed dataset. The three-LLM-judge scoring method is also worth watching: it's a reasonable stand-in for human relevance judgment, but it's not the same as conversion data or actual shopper behavior, and Onton picked the judges. The exclusion of image queries from the main 90, for defensible reasons tied to Lens limitations, still means the harder multimodal case gets a separate, smaller comparison.

For founders building search or recommendation products, the real signal here isn't the 2.7x headline. It's whether neurosymbolic retrieval holds its accuracy as catalog size scales toward Amazon's actual inventory. That test hasn't been run yet, at least not publicly.

Common Questions Answered

What is Ontology 1 and how does it perform compared to Google Shopping and Amazon search?

Ontology 1 is a neurosymbolic model developed by Onton for complex, conversational, and multimodal product search. On a 90-query benchmark, it achieved a mean precision@10 of 0.630, outperforming Google Shopping's 0.543 and Amazon's 0.469, representing approximately 2.7x greater accuracy than Amazon's search engine.

Why is Onton's achievement of indexing only 1% of catalog size significant?

Onton's ability to achieve higher precision@10 scores while indexing only about 1% of the catalog that Amazon and Google Shopping use suggests a fundamentally different approach rather than relying on brute-force indexing of massive datasets. This indicates that Ontology 1 may use a more efficient retrieval architecture or better-curated dataset strategy rather than simply processing more product data.

What methodology did Onton use to score Ontology 1's accuracy in their benchmark?

Onton's benchmark consisted of 90 queries that were evaluated by three independent LLM judges to determine precision@10 scores. This multi-judge approach was used to assess how accurately Ontology 1 retrieved relevant products compared to competitors.

What potential limitations exist in Onton's benchmark results?

The benchmark tested Ontology 1 against a much smaller catalog (1% of competitors' sizes), which may make precision measurements easier to achieve with a pre-curated dataset rather than proving superior retrieval architecture. Independent verification at larger catalog coverage levels (10% or 50%) would be needed to definitively prove Ontology 1 has fundamentally better search technology rather than simply better-groomed data.

LIVE02:59Onton Claims 2.7x More Accurate Search Than Top E-commerce Engines