Editorial illustration for DesignArena creators raise USD 7.9 million for AI taste-testing platform
DesignArena Raises $7.9M for AI Taste-Testing Platform
Grace Li and her college friends set out to build an AI game engine in early 2025. The models worked fine, technically. The games they produced just weren't fun, and nobody on the team could explain why in a way an algorithm could use.
That problem, judging taste rather than function, turned into a business. The tool they built to solve it, DesignArena, now has 5.3 million users testing AI-generated designs against each other in head-to-head matchups. On Monday, the company behind it, called Intelligence, announced a $7.9 million seed round led by Index Ventures, with Conviction's Sarah Guo and Mike Vernal, A*, and Valkyrie also putting in money.
The pitch is simple: AI labs can generate infinite variations of a website or an image, but they have no reliable way to know which ones people actually like. DesignArena sells them that answer, one comparison at a time. Li says the company's first major deal with a frontier lab came together within a week of realizing what they'd stumbled into.
It’s a useful service, but the real value of the platform comes from the enterprise side, where participating models can treat it as a source of endless instant feedback for their media-generating models. The users tend to be indifferent to which models they’re ranking — as Li puts it, they just want the best output they can get — so their rankings can give critical input to what users really want.
Why this matters
DesignArena's pitch works because it turns an old problem, taste, into a data pipeline. Founders like Grace Li stumbled into this from a game engine that produced technically correct but boring output, then realized the fix wasn't more compute, it was human judgment collected at scale. That's the part worth watching.
Labs already lean on RLHF and benchmark leaderboards, but those measure narrow correctness, not whether something is actually good to look at, play, or read. A $7.9 million round suggests investors think aesthetic ranking is becoming its own infrastructure layer, not a nice-to-have.
For developers and founders building media-generating models, this is a signal that user preference data is turning into a product category, one you might soon rent instead of build. For researchers, it raises a real question about what "indifferent" users ranking anonymous outputs actually teaches a model about taste versus popularity. We'd want to see how DesignArena handles bias in who's doing the ranking before treating its scores as ground truth. Worth tracking whether enterprise clients disclose how they weight this feedback against their own internal evals.
Common Questions Answered
How does DesignArena help AI models improve their media-generating capabilities?
DesignArena provides enterprise clients with a source of endless instant feedback by collecting user rankings of AI-generated designs in head-to-head matchups. This user preference data serves as critical input that helps models understand what users actually want, going beyond traditional narrow correctness metrics like RLHF and benchmark leaderboards.
What problem did Grace Li and her team originally encounter when building their AI game engine?
While the AI models worked fine technically, the games they produced were not fun and the team could not explain why in a way an algorithm could understand. This challenge of judging subjective taste rather than technical function led them to create DesignArena as a solution.
How many users currently participate in testing designs on DesignArena?
DesignArena has 5.3 million users who test AI-generated designs against each other in head-to-head matchups. These users are primarily focused on getting the best output possible rather than caring about which specific models they are ranking.
Why is DesignArena's approach to collecting human judgment at scale significant for AI development?
DesignArena converts the subjective problem of taste into a scalable data pipeline that measures whether something is actually good to look at, play, or read—not just whether it is technically correct. This represents a departure from existing methods like RLHF and benchmark leaderboards, which focus on narrow correctness metrics rather than subjective quality.
Further Reading
- Design Arena: World's largest crowdsourced benchmark for AI-generated design - Y Combinator
- Design Arena: Funding, Team & Investors - Startup Intros
- Launch HN: Design Arena (YC S25) – Head-to-head AI design benchmark - Hacker News
- Design Arena - AI Design Model Benchmark - EveryDev
- Design Arena: The Public AI Benchmark for Smarter Choices - TechPilot