AI · SEMICONDUCTOR · FACT CHECK

SKD WIRE

EN WIRE · 2026-10-03 12:11

Accounting Task AI Scores '100%' vs Humans '37%' — Reversal Comes 18 Months After AI Trailed at 'Perfect' Benchmark Stage

Mostly True

Accounting Task AI Scores '100%' vs Humans '37%' — Reversal Comes 18 Months After AI Trailed at 'Perfect' Benchmark Stage
Image source: 조사 출처

An AI system has achieved a perfect 100% score on an accounting task benchmark, where human professionals scored 37%, marking a reversal roughly 18 months in the making as the technology reached what the original assessment described as the "perfect" stage.

원문 주장 (KR)
회계 업무서 AI '100%' vs 인간 '37%'...역전 18개월 만에 '완벽' 단계로

From Behind to Front

The comparison shows AI outperforming humans on accounting work by a wide margin — a full 63 percentage-point gap. The milestone is notable because, according to the claim, the positions were reversed 18 months earlier, when the benchmark stood at a stage where AI had not yet reached the "perfect" level.

Benchmark Context

Accounting tasks have been viewed as a demanding test for large language models (LLMs), requiring precision and adherence to rules rather than open-ended generation. A 100% result against a 37% human score, if it holds, would represent one of the clearest cases of AI surpassing professionals on a structured knowledge task. The precise identity of the benchmark, the group of human test-takers, and the testing conditions were not detailed in the available material.

What Remains Unclear

The available information does not specify which AI model produced the 100% score, nor how the human participants were selected or measured. Details such as the task format, scoring methodology, and whether the human baseline reflects experienced accountants or a broader sample remain unconfirmed. Readers should treat the comparison as accurate in its headline figures but limited in verifiable context.

Bottom Line

The core numbers — AI at 100% and humans at 37% on the accounting task, roughly 18 months after the technology had not yet reached that level — match the claim as presented, though supporting details are sparse. On that basis, the assessment is Mostly True.

Sources — primary documents (8)
  1. https://www.aitimes.com/news/articleView.html?idxno=215935
  2. https://www.mercor.com/blog/human-baselines-for-benchmarks-ai-now-outperforms-junior-accountants/
  3. https://cdn.sanity.io/files/h6s14f4z/production/4048fb3eece18a56006026d212d25312f7ffc29a.pdf
  4. https://www.mercor.com/apex/apex-accounting-leaderboard/
  5. https://cellcog.ai/blog/ai-vs-accountants-mercor-study/
  6. https://cdn.sanity.io/images/h6s14f4z/production/4154d9753c98533da19506b9954b9a8dc650440b-2200x1062.png
  7. https://cdn.sanity.io/images/h6s14f4z/production/0ec454bcb2fba47e59b51bd9bd2ebcd2799dee58-2320x1660.png
  8. https://cdn.sanity.io/images/h6s14f4z/production/46639e0832dc06d155e05c557a74b3d5f63a3fda-2320x988.png

KR: /news/20261003-88a5e1 · 판정: 대체로 사실