Editorial illustration for Optima's New AI Benchmark Lets Users Test Models With Their Own Data
Optima AI Benchmark: Test Models With Your Data
Artificial Analysis, the research group behind independent LLM evaluations and benchmarks like GDPval-AA and AA-Briefcase, has launched a new platform called Optima. The pitch is straightforward: standard AI benchmarks measure models against fixed tasks that rarely match what a company actually needs done. A model that scores well on a public leaderboard might still be the wrong pick for, say, processing insurance claims or summarizing legal contracts. Optima lets users skip that mismatch entirely by building benchmarks from their own data and workflows.
Instead of relying on someone else's test set, users feed Optima their own inputs, whether that's raw data, a description of a task, or sample outputs showing what "good" looks like. The platform then runs current leading models against that material and reports back on quality, cost per task, and time per task. That last part matters for anyone weighing a cheaper, slower model against a pricier, faster one for a specific job. Optima is live now, and Artificial Analysis frames it as a way to answer a question generic benchmarks can't: which model actually works best for this particular problem.
Users can build their own benchmarks using their own data, workflows, or descriptions of a use case, then run them across leading current models and compare results on quality, cost per task, and time per task, according to Artificial Analysis.
Why this matters
Public leaderboards have always told us how a model performs on someone else's data, which is a different question from how it'll perform on yours. Optima's bet is that teams building actual products care more about cost per task and latency on their own workflows than a percentile rank on MMLU or whatever benchmark is trending that month. That's a reasonable bet, and it maps onto a complaint we hear constantly from engineering teams: the model that tops the chart is often not the model that works for their specific document format, customer queries, or pricing constraints.
The open question is rigor. Custom benchmarks are only as good as the sample inputs and outputs users feed in, and a platform that lets anyone spin up a test also lets anyone spin up a bad one. If Artificial Analysis can keep Optima's results comparable and resistant to cherry-picking, this becomes a genuinely useful procurement tool for teams choosing between providers. If not, it just adds another number to argue about.
Common Questions Answered
What is Optima and how does it differ from standard AI benchmarks like MMLU?
Optima is a new platform launched by Artificial Analysis that allows users to create custom benchmarks using their own data and workflows instead of relying on fixed public tasks. While standard benchmarks measure models against generic datasets that may not reflect real-world use cases, Optima lets companies test AI models on tasks specific to their needs, such as processing insurance claims or summarizing legal contracts.
What metrics can users compare when testing models on Optima?
Users can compare leading AI models across three key metrics: quality of results, cost per task, and time per task. This allows teams to evaluate which model best fits their specific requirements and constraints, rather than relying solely on leaderboard rankings.
Why does Optima address a major flaw in current AI benchmarking practices?
Public leaderboards measure how models perform on someone else's data, which often doesn't predict performance on a company's actual workflows and data. Optima solves this problem by enabling teams to test models against their own use cases, ensuring they select the model that truly performs best for their specific business needs rather than one that simply ranks high on trending benchmarks.
Who created Optima and what is their background in AI evaluation?
Optima was launched by Artificial Analysis, a research group known for creating independent LLM evaluations and benchmarks including GDPval-AA and AA-Briefcase. The team has established expertise in evaluating and comparing language models across various metrics.
Further Reading
- Announcing Optima: create a custom benchmark for your use case - Artificial Analysis
- Optima - Artificial Analysis - Artificial Analysis
- Artificial Analysis Intelligence Index v4.1: a shift toward ... - Artificial Analysis
- Launching v4.1.1 of the Artificial Analysis Intelligence Index - Artificial Analysis
- Artificial Analysis launches Optima: private evals from your own traces - ThursdAI