undefined New: Digits adds 11 new quality checks to make closing the books that much easier! Learn more

Bar chart comparing transaction categorization accuracy. Digits AGL® leads at 97.8%, followed by OpenAI GPT-6 Luna at 79.5%, Google Gemini 3.8 Flash at 79.2%, human outsourced accountants at 79.1%, OpenAI GPT-6.1 Sol at 78.9%, Anthropic Claude Fable 5.1 at 78.2%, Anthropic Claude Opus 5.5 at 77.9%, and TypeSafe Jev at 61.4%.

Breaking: AI Can’t Close Your Books From the Outside (Yes, Including Jev)

In June, we reported that frontier models had caught up with outsourced human accountants at categorizing transactions. Since then, nearly every major AI lab has shipped a new model with powerful new capabilities, and TypeSafe launched Jev, a much-hyped “deterministic” model, so we got very curious… would these new AI models finally be good at accounting? 

Our machine learning team diligently re-ran our Beyond the AI Hype benchmarks, and the answer is: No.

The short version: the models didn’t get any better at accounting on their own. Not with newer models, not with more expensive ones, and not with more reasoning. Jev got a lot of hype. It’s a brilliant model for some things, but it’s terrible at accounting. One-shot LLM classification has plateaued (and actually declined since June). They did get better when they could look up how the business had booked things in the past, with the tradeoff of extreme cost and token usage. It remains the case that ledger-native, domain-specific accounting models such as Digits AGL® continue to dramatically outperform all frontier labs. 

The three approaches in AI accounting

We compared three ways of getting AI to do your books. In each case, the AI had to categorize 2,000 real transactions from four small businesses, and each answer was graded against categories reviewed by U.S. accountants applying GAAP principles:

  1. A general-purpose AI with no context. You hand it a transaction and the chart of accounts, then ask it to pick a category using general accounting knowledge. This is basically what happens when you drag bank statements into Claude or ChatGPT.

  2. The same AI with access to the ledger. You give it a tool to look up how the business booked similar transactions in the past (plus web search).

  3. Ledger-native agents. AI agents that live inside the general ledger and learn from each business's own books. The one we studied was Digits AGL®, our Agentic General Ledger.

We tested 84 configurations: the latest models from OpenAI, Anthropic, Google, DeepSeek, Moonshot AI, Z.ai, Meta, and TypeSafe, at every reasoning level each one offers.

No context: the models have plateaued

If you've wondered whether ChatGPT can do your bookkeeping, this is the test that answers it. The best model working alone was OpenAI's GPT-6 Luna at 79.5%, which is slightly lower than June's best (80.7%). The next six models all finished within half a point, so they're effectively tied. And they're about as accurate as a human accountant seeing these books for the first time (79.1%).

Turning up the reasoning didn't help much either. No amount of thinking will tell a model how a client books things when it has never seen that client's books.

With access to the ledger: better, but not enough

When the same models could check the books, accuracy jumped. The best agent, Claude Fable 5.1, reached 88.9%, about 10 points higher than the model scored in a one-shot configuration.

The gain came from looking things up. For example, one business books its home-improvement store purchases to Fulfillment because the materials go into what it sells. Every model working without context got all 57 of those transactions wrong. The agent checked the history, found 15 earlier purchases booked to Fulfillment, and got it right. Web search barely mattered, because it can tell you what a merchant sells but not how this business books it.

Think of the model on its own as a smart new hire on day one: they know accounting, but they don't know the client yet. The agent is that same new hire after they've checked last year's books.

But that’s still not enough:

  • It's still wrong about one time in nine. Half of the best agent's misses happened because it never checked the right category.

  • It's slow and expensive. The best agent took 74 seconds and 16,295 tokens per transaction. Multiply that by every transaction, every client and every close.

  • It doesn't always finish. One model left 20.7% of transactions with no answer, because it was still calling tools when it ran out of turns.

  • It misses house rules. Many of the remaining errors involve a business's own conventions, like overlapping categories, payment processors that hide the real vendor, and one-off quirks. The top three agents come from different providers, yet they picked the same wrong category on 163 transactions.

Native to the ledger: Digits AGL® results

Digits AGL® classified 97.8% of transactions correctly. That is 18.3 points more accurate than the best one-shot model and 8.9 points more accurate than the best agent. Measured by error rate, AGL® got 2.2% of transactions wrong compared with 11.1% for the best agent: one-fifth as many errors.

Digits did this in 40 milliseconds and 64 tokens per transaction, which is 1,846 times faster and 255 times fewer tokens than the best agent configuration.

Of the transactions that all three top agents got wrong, Digits AGL® classified 83% correctly, which tells us these errors can be learned from the business's own history. An agent rediscovers that history with a few lookups per transaction, and only for the categories it thinks to check. AGL® is trained on each business's ledger, so the conventions are already part of how it classifies.

What this means for accounting: ledger-native AI wins

Everyone can rent the same frontier models, so the model itself is no longer the advantage. But this benchmark shows that access to context isn’t the advantage either. The best agents could look up the same business's history and still finish nearly nine points behind ledger-native agents. The difference is whether context is looked up or learned. 

An agent with ledger access is like a student taking an open-book exam against the clock. It finds the answer in a textbook when it knows which page to turn to, and misses when it doesn’t. A ledger-native agent has already learned the materials so it’s more accurate, faster, and costs a fraction as much to run every month.

The full study, including every configuration's results, the failure analysis, and the agent trace analysis, is in the whitepaper.

Download the whitepaper

Switch to Digits today

Experience accounting, reimagined.

Abacus icon

Accounting Firms

Build an AI-native practice with Digits.

Get started

For firms of every size

Building icon

Businesses

Automate bookkeeping for your small business or startup with 24/7 AI.

Get started

Free 30-day trial