Claude Opus 5 Scored 100 Percent on Accounting Tasks Where Licensed CPAs Averaged 37 Percent

Mercor recruited 12 licensed CPAs and gave them four month-end close scenarios drawn from its APEX-Accounting benchmark. Each scenario required the accountant to dig through a company's working files, find the right figures, do the math, and deliver a table of results. The participants averaged about five and a half years of experience and worked without AI assistance. Their accuracy ranged from 0 percent to about 90 percent, the average was about 37 percent, and most attempts took between 30 and 180 minutes.
Claude Opus 5 ran the same tasks 20 times. It scored 100 percent on all 20 attempts and finished each one in under 10 minutes. Mercor researcher Aden Barton, who wrote up the study, reports that the model beat even the best human participant on both accuracy and time.
How the study was built
The tasks were simplified versions of items from APEX-Accounting, a benchmark Mercor built to stump frontier models. Because of that origin, they contain realistic but hard-to-spot requirements that compound on one another, so a single overlooked number can drag a score very low. Mercor's own task authors expected averages of 30 percent for a junior accountant and 55 percent for a mid-level one. The CPAs landed at about 37 percent, close to the junior expectation.
Mercor says it originally designed the study to measure augmentation, meaning how much an accountant improves with AI assistance. That arm was dropped. The model alone hit the ceiling, which left no room to measure uplift, so the comparison reduced to unassisted humans against the model.
A fast climb
Mercor plotted model scores on these tasks by release date. Eighteen months ago the best models fell below the 37 percent human average, and GPT-4o scored near zero. OpenAI's o3 passed the average in spring 2025, GPT-5 reached about 69 percent, and Opus 5 and other recent frontier models now sit at or near 100 percent. Some budget and open-weight models, including Qwen3.5-122B, still score below the human line.
Cost followed the same direction. Mercor puts Opus 5 at $0.21 per rubric criterion met and the accountants at $10.35, a ratio of 49 to 1. The human figure is built from the US median accountant wage, so it reflects what the time would cost at typical pay rather than what any particular firm pays. The model figure does not appear to include the review time a firm would add before relying on the output, although the blog post does not say either way.
What Mercor says the results leave out
The authors are careful about scope, and some of their caveats are substantial. The tasks reward detail-oriented instruction following and searching through files, which are skills where models are strongest. They do not cover client communication, asking the right clarifying questions, or building up tacit knowledge of a company over months.
The setting also removed supports that accountants normally have. Participants had no coworkers to ask and no accumulated job context, which made the tasks harder for them than the same work would be in an office. Mercor says it considered not publishing the results at all because of the risk of misreading, and chose to release them for the sake of transparency.
There is also a point about how benchmarks evolve. Tasks get made harder to induce model failures, in the way exam questions are written to challenge top students. As models improve, some tasks take many experts working together for tens of hours to build, which means a single person could not complete them in a normal setting. Mercor argues that benchmarks are shifting toward work only AI can fully do, and that this shift makes sense given where economic value is likely to sit. Its post opens with the observation that benchmark scores have soared while AI adoption shows up only faintly, if at all, in US productivity statistics, and it presents the human baseline as one way to connect the two.
The longer benchmark tells a different story
The 100 percent result applies to medium-length, well-defined tasks. Mercor's full APEX-Accounting leaderboard measures long-horizon work that crosses accounting software, spreadsheets, and PDFs. On that leaderboard the top entry is Claude Opus 5.5 at 61.8 percent, followed closely by Fable 5.1 at 61.0 percent. Mercor's announcement of the Opus 5.5 result reports a Pass@1 of 15.4 percent on APEX-Accounting, meaning the model produced a fully passing solution on its first attempt for roughly one task in seven. The same post says Opus 5.5 passed 142 of 239 tasks on all four runs in the broader APEX-Agents benchmark, which covers several professions rather than accounting alone.
Taken together, the two Mercor datasets describe a model that is flawless on bounded close steps and still inconsistent on the multi-application close that a finance team actually runs at month end. A secondary summary of the study from an AI news digest draws the same line, saying the useful split is task length rather than job title. That reading fits the numbers, though it is a commentator's interpretation and Mercor does not frame it that way itself.
Interpreting a vendor study
Several features of the study call for care before the headline is repeated. The sample is 12 accountants on four tasks, with Opus 5 contributing 20 attempts. The accountants were recruited and tasked by a company that sells human data and evaluation services and that built the benchmark in question. The full paper sits behind a download link that automated tools cannot open, so this article relies on Mercor's blog summary and press coverage of it rather than the paper's tables. The study has not been peer reviewed.
The headline gap is also partly a product of design. The tasks were tuned to catch models out and then given to humans stripped of normal workplace support. A model that aces them has shown it can follow dense, trap-laden instructions across a file set. That result is meaningful for the structured parts of a close and says little about the parts that depend on judgment and conversation.
What firms can do with this
Accounting teams weighing these results can run their own baselines. A short exercise using their staff, their files, and a rubric drawn from real close tasks would show whether a vendor's benchmark gap holds in their environment. Cost comparisons should include the time spent reviewing model output and the cost of failures on long, cross-application work, since the 49-to-1 ratio covers neither. Firms evaluating vendors can ask for the human baseline behind any accuracy claim, along with the task length and the number of applications involved.
Mercor's blog post and the linked paper are the starting points for anyone who wants the primary numbers, and the APEX leaderboard is updated as new models are scored.

