Home About Shannon Services Industries San Diego Blog Tools & Resources Contact Book Now

Producing a convincing financial record used to take effort, and that effort was quietly doing protective work nobody accounted for. AI has removed it from both sides of your desk at once, which changes what your books need from you and what the people trying to steal from you can now afford to build.

A benchmark published in June handed AI models the same pile of paperwork you would hand a bookkeeper. Invoices, bills, contracts, an opening trial balance, and a few documents that did not belong. The task was to produce journal entries with the source document cited for each one.

The best model produced an exactly correct balance sheet 46% of the time. Under strict scoring, where every entry had to trace to the right document, that fell to about 5%.

The finding that matters most is this one: in 52% of the entries where the model got the accounts and the amounts right, it cited the wrong supporting document (FinBalance, June 2026).

That produces a set of books where the entries balance, the categories look sensible, and the paper trail points at the wrong paper. Books that look documented right up until someone pulls a source.

What AI Is Actually Good At Now

The tools your business already runs on have shipped real automation in the past eighteen months. QuickBooks Online now has eight AI agents, including one that categorizes transactions and asks clarifying questions when something is ambiguous. Dext launched AI Assist in March, which learns a firm's own categorization decisions and applies them consistently across clients. BILL codes multi-line invoices from historical patterns. Xero pulls itemized detail out of a document in about thirty seconds. None of that fixes what sits underneath it, and a tool categorizing confidently on top of a messy chart of accounts makes the common QuickBooks Online setup mistakes harder to see, not easier.

Capture, extraction, first-pass categorization, straightforward transaction matching. The parts of this work that were volume and repetition are being handled competently by software right now, and the Bureau of Labor Statistics projects bookkeeping clerk employment to fall 6% through 2034, attributing the decline directly to that automation (BLS).

If a bookkeeper's value was being faster at data entry than you were, that value is expiring. Data entry was never the job.

Where It Fails

Two findings from that benchmark matter more than the headline accuracy numbers.

The first is that the failures do not degrade gradually. On the simplest document bundles, models produced a correct balance sheet 88% of the time. On the most complex, 0%. The easy work is nearly solved, the hard work is not partially solved, and the drop between them is steep rather than gentle.

The second finding is stranger. When the researchers let the models check their work against a ledger, balance sheet accuracy improved substantially, and their ability to notice that a bundle was internally inconsistent got worse. By as much as 52 percentage points on the smaller models. Making the books balance made the system worse at noticing when they should not have balanced at all.

That gap is the job. Reconciling is not arriving at a matching number. It is the judgment that the number should not match yet, that a document is missing, that this invoice contradicts that contract.

Two honest caveats about the study. It used synthetic documents with clean text and one fixed chart of accounts, and its authors say plainly that it measures reasoning over tidy inputs rather than real-world document handling. A narrowly tuned commercial product may well do better on its own turf.

How much better is not something anyone outside the vendors can currently answer. Every accuracy figure in this market is vendor-published, measured against test sets nobody else can see. BILL reports 99% accuracy extracting key invoice fields and 92% accuracy on transactions processed entirely by AI (BILL). Both figures can be true, because they measure different things, and the second one is the one that matters to you: roughly one transaction in twelve is wrong. There is no independent evaluation of any of these products running on real client books.

The Thing Nobody Priced In

Producing a plausible financial record used to take effort. Someone typed the entry. Faking a receipt convincingly required Photoshop skills and patience. Verifying a vendor's bank change required a phone call, which cost five minutes, which is exactly why people made the call.

None of that friction was designed as a control. But effort is a filter, and it was doing protective work that never appeared on anyone's org chart.

AI has collapsed that cost to nothing, on both sides of your desk. Your software produces a coded, finished-looking transaction in seconds, and so does the person trying to steal from you. The two outputs look equally polished because they came from the same class of tool. Producing the record used to be the scarce thing. Now verifying it is.

What the Fraud Actually Looks Like

Fake Receipts Are the Measured Shift

AppZen screens expense reports for its clients. Of the receipts it flagged as fraudulent, none were AI-generated in March 2025. By mid-May 2026, 70.8% were (Accounting Today). Traditional template fakes fell from nearly all of its catches to 29%. That figure describes one platform's flagged pile rather than all expense fraud everywhere, and it has been widely repeated as the latter. What is meaningful is the slope: fourteen months from zero to the majority.

The dollar value is the part that should change your controls. AI-generated fakes averaged about $100, against $182 for the older template versions. The new fraud is deliberately smaller, sized to sit under the auto-approval threshold where nothing routes to a human for review. And it is not coming from a distant criminal. In a survey of 2,000 US and UK workers, 34% admitted generating a fake receipt with AI.

Synthetic Vendors Are the Sophisticated Version

The Journal of Accountancy reported in July on criminals assembling complete fake vendor packages, including generated W-9s, fabricated tax documents, and a professional website. Some of them staff the fake vendor's phone number with an AI voice agent that sounds like a real person (Journal of Accountancy).

That defeats the control most businesses believe they have. If your process is to call the vendor and verify, and the number you are calling came off the invoice, you now reach a bot that confirms everything. It is the same playbook as the scams already aimed at business owners, with the manual labor taken out of it.

The Deepfakes Are Real but Oversold

The FBI's Internet Crime Complaint Center logged $3.05 billion in business email compromise losses in 2025, and attributes just over $30 million of that to an AI nexus (IC3 2025 Annual Report). About one percent. There is real undercounting in that number, since a victim usually cannot tell a cloned voice from a real one, but the gap between one percent and the volume of coverage is worth saying out loud.

Meanwhile the unglamorous version is thriving. 74% of organizations were hit by business email compromise in 2025, up from 63% the year before, and checks remain the most-targeted payment method at 58% (AFP 2026 Payments Fraud Survey). What is coming for your business is a well-written email, a clean fake receipt, and a vendor record nobody verified.

What Still Works

Call Out, Never In

JPMorgan published guidance in February that is unusually specific about how callback verification fails. Do not let a vendor call you to confirm instructions, and never use a phone number that came from the email or document itself. Call out, to a number from your own records, and speak to the person actually responsible for the change (JPMorgan). With AI voice agents now answering vendor lines, where that number came from is the entire control.

Verify Knowledge, Not Appearance

In 2024, a Ferrari executive received calls from what sounded exactly like the company's CEO. He asked what book the CEO had recommended to him days earlier. The call ended. When impersonation becomes effectively perfect, "does this sound right" stops being a test, and "does this person know something only they would know" still works. The FBI now recommends agreed verification phrases for the same reason.

Stop Trusting Your Thresholds

If fraud is being sized to fit underneath your auto-approval limit, then a static dollar threshold has quietly become a targeting signal rather than a protection. That does not mean abandoning thresholds. It means sampling below them, and reviewing the vendor list and payroll register monthly instead of only watching the totals.

Document How AI Touched the Work

For anyone working in tax, this stopped being optional in June. The IRS Office of Professional Responsibility issued a bulletin applying Circular 230 to AI use, requiring practitioners to verify the accuracy of the facts, citations, and calculations that AI produces, and to treat its output as a starting point rather than a finished product (The Tax Adviser). Given a 52% wrong-citation rate in the benchmark above, verifying citations is not a formality.

The Bottom Line

The businesses that catch fraud early are the ones where two people's records have to match. I have written before about separating bookkeeping from bill-pay, and the pattern holds here: almost every one of them depends on a single person being able to create a transaction and approve it without anyone else seeing it.

AI does not change that principle. It changes the price of ignoring it. A sloppy process used to survive on the sheer difficulty of producing a convincing document, and it will not now. Not against someone building a vendor out of nothing, and not against your own books quietly citing the wrong invoice for six months.

So no, I do not think AI is coming for this profession. It is taking the part of the job that was never really the job, and leaving the part that was always harder to hire for: knowing what should be there, and being willing to ask a question the software cannot.

See how we approach clean, reviewed books: Our Bookkeeping Services

Questions about your books? Reach us at info@saltandsandbookkeeping.com or (619) 304-SALT (7258).

Not sure who is actually reviewing what in your books? Let's find out.

Schedule a Free Consultation
← Back to all posts