Editorial illustration for AI solves accounting tasks flawlessly in benchmark with 160 tasks
AI Beats CPAs on 160 Accounting Tasks in New Benchmark
AI solves accounting tasks flawlessly in benchmark with 160 tasks
Twelve licensed CPAs sat down with a simplified set of bookkeeping tasks pulled from the APEX Accounting Benchmark. On average, they had five and a half years of experience each. The company running the test, Mercor, wanted to see how they'd stack up against AI models on the same work.
Eighteen months ago, that comparison wasn't close, and the humans won. The best AI models at the time scored below the accountants' average of about 37 percent.
That gap has closed fast. The full APEX Accounting benchmark now runs 160 tasks across 10 simulated companies, built by a pool of more than 40 professionals averaging 11 years in the field. It's a much tougher test than the simplified version the CPAs took, covering the kind of detail-chasing, instruction-following work that bookkeeping is full of. Claude Opus 5.5, Fable 5.1, and GPT-6 Astra are the current front-runners on this harder benchmark, and the scores show just how much ground AI has made up in a year and a half, even if none of them are anywhere near solving every task outright.
Mercor also admits the study's tasks test exactly what AI does best, which is hunting down details and following instructions precisely. The study left out key parts of the job, such as talking with clients, checking in with colleagues, and drawing on context built up over years. Mercor says that's why accountants can't be replaced, though it expects major productivity gains across the industry.
Why this matters
The gap between the two numbers here is the real story. Models go from "almost flawless" on simplified tasks to 61.8 percent on the full 160-task benchmark, built by accountants with 11 years of experience on average. That's progress, but it's also a reminder that benchmark headlines about AI "beating" professionals often measure the easy version of the job.
For founders building bookkeeping or finance-automation products, this is the window to move: speed and cost advantages are real and measurable now, not theoretical. For researchers, the 18-month jump from below 37 percent to near-perfect on simplified tasks says more about how fast models improve on structured, rule-based work than about general competence. Accounting has clear rules and checkable outputs, which is exactly where current models do well.
Messier judgment calls, the ones that make up the rest of that 160-task set, are where the score drops. Read this as license to automate the routine parts of bookkeeping, not the sign-off. Someone with a CPA still needs to own the final number.
Common Questions Answered
How did AI models perform on the APEX Accounting Benchmark compared to licensed CPAs?
AI models achieved nearly flawless performance on the simplified bookkeeping tasks, significantly outperforming the twelve licensed CPAs who averaged about 37 percent accuracy. However, on the full 160-task APEX Accounting Benchmark, AI performance dropped to 61.8 percent, demonstrating that while AI excels at simplified tasks, it struggles with the complete complexity of real accounting work.
What specific accounting tasks does the APEX Accounting Benchmark test?
The APEX Accounting Benchmark consists of 160 tasks designed by accountants with an average of 11 years of experience. The benchmark tests AI's ability to hunt down details and follow instructions precisely, which are areas where AI performs exceptionally well, but it also includes complex real-world accounting work that requires contextual understanding.
Why can't AI completely replace accountants according to Mercor's findings?
Mercor acknowledges that the benchmark tasks test only what AI does best, excluding critical aspects of accounting work such as client communication, colleague collaboration, and drawing on context built up over years of experience. These human-centric elements of the job cannot be easily replicated by AI, which is why accountants cannot be fully replaced despite AI's superior performance on technical tasks.
What productivity gains does Mercor expect AI to bring to the accounting industry?
While Mercor stops short of providing specific productivity percentages, the company expects major productivity gains across the accounting industry as AI takes over the technical, detail-oriented tasks where it performs nearly flawlessly. This suggests that accountants will be able to focus more on client relationships and strategic work while AI handles routine bookkeeping and data processing.
How much has the performance gap between AI and accountants changed in the past 18 months?
Eighteen months before the study, the best AI models scored significantly below the accountants' average of about 37 percent on accounting tasks. By the time of the benchmark test, AI models had closed this gap dramatically, achieving nearly flawless performance on simplified tasks, though still falling short at 61.8 percent on the full 160-task benchmark.