# Can You Trust AI to Categorise Transactions Correctly?

> Can you trust AI to categorise transactions? Yes, with a review step. Accuracy usually passes 90 percent after training, but the last few percent are where it matters most. Here is how AI categorisation works, where it reliably goes wrong, and how to build a workflow you can actually trust.

**Source:** https://agentive.au/blog/can-you-trust-ai-to-categorise-transactions/ · **Published:** 2026-08-18 · **Author:** Dr. Ash Khalilian

---

Short Answer

**Yes, with a review step.** AI is accurate enough to draft your transaction coding, not to sign it off unseen. Accuracy typically passes 90 percent once it has learned your patterns, but the remaining few percent are the ones that matter: new vendors, ambiguous or mixed transactions, one-offs, and GST edge cases. The trustworthy model is simple. AI drafts, a human reviews the exceptions, and every correction makes it better. Trust comes from that process, not from faith in the machine.

Scroll through any accounting forum and you will find the same two threads on repeat. "Does anyone actually like the AI in their accounting software?" and "How are you using AI in your practice?" Underneath the venting and the curiosity is one practical worry that nobody says out loud: if I let this thing code my transactions, can I actually trust what comes out? Nobody wants to explain a botched BAS because a bot quietly filed a capital purchase as office snacks.

It is a fair worry, and the honest answer is more useful than a yes or a no. AI transaction categorisation is genuinely good, good enough that roughly 98 percent of accountants and bookkeepers reported using AI in some form over the past year. But good is not the same as trustworthy without a process around it. This article explains how the categorisation actually works, why the headline accuracy number can mislead you, exactly where it tends to go wrong, and how to build a workflow you can rely on. For the wider view, we also covered [whether AI will replace bookkeepers and accountants](/blog/will-ai-replace-accountants-bookkeepers/) separately.

## How AI Actually Categorises a Transaction

It helps to know what is happening under the bonnet, because it demystifies both the strengths and the failure points. AI categorisation is pattern matching, not magic. For each transaction it reads the available signals, weighs them against what it has seen before, and proposes a code. Here is what it looks at.

### The Transaction Details

Payee name, amount, date, bank description, and any reference. A payment to a known supplier for a familiar amount is an easy call.

### Your Coding History

How similar transactions were coded before. If you have always put this vendor to subscriptions, it learns to do the same.

### Linked Documents

The matched invoice or receipt, when one exists, gives line-level detail and the GST shown on a valid tax invoice.

### Its Own Confidence

Good systems score how sure they are and surface low-confidence guesses for review, rather than posting them silently.

The key insight is that last box. A categorisation engine that quietly posts everything at the same confidence is far riskier than one that says, in effect, "I am 96 percent sure about these 40 and genuinely unsure about these 3, please look." That distinction is the whole difference between a tool you can trust and a tool that surprises you at BAS time.

## Why 90 Percent Accuracy Is Both True and Misleading

You will see the same figure everywhere, and it is roughly right. Once a system has learned a firm's patterns, transaction categorisation accuracy typically exceeds 90 percent. That sounds reassuring, and for the bulk of your ledger it should be. The problem is what the average hides.

The Number That Matters

Accuracy above **90 percent** is not spread evenly across your transactions. The easy, recurring ones are close to perfect. The hard ones cluster all the errors. So the last few percent are not random noise you can shrug off. They are disproportionately the unusual, high-stakes transactions that actually move your numbers and your GST position. Which is exactly why a human still has to look.

Think of it this way. If ninety-something percent of your transactions are routine coffee, software subscriptions, and regular suppliers, an engine that nails all of those and misses a handful of tricky ones will still score beautifully on paper. But the ones it missed might be the equipment purchase, the mixed personal-and-business expense, or the GST-free item you now have to unpick. The average looks great. The exceptions are where the risk lives, and reviewing them is how the trust gets earned.

## Where AI Categorisation Reliably Goes Wrong

The good news about the errors is that they are predictable. AI does not go wrong randomly, it goes wrong in the same places every time. Once you know the failure modes, a review stops being a slog through everything and becomes a targeted check of the usual suspects. These are the five that come up again and again.

### New Vendors It Has Never Seen

With no history to learn from, a first-time supplier is a genuine guess. The AI will lean on the name and amount, which is fine for an obvious utility bill and shaky for a payee whose name gives nothing away. New vendors deserve a first-time human check, after which the pattern is set.

### Ambiguous Transactions

A payment to a large retailer or a general marketplace could be office supplies, equipment, or stock. The payee alone does not tell you the purpose, and the AI cannot read your intent. These need the context only the person who made the purchase, or who knows the business, can supply.

### Mixed and One-Off Transactions

A single payment that should be split across two accounts, or a genuine one-off with no pattern to learn from, are both hard for pattern matching by definition. The AI will usually pick the single most likely code, which is wrong when the honest answer is "this needs splitting" or "this is unlike anything before it".

### GST Treatment Edge Cases

This is the one Australian bookkeepers watch closely. GST-free items like certain food and health lines, mixed-supply receipts where only part is taxable, and purchases with no valid tax invoice all trip up automated coding. The account might be right while the GST is wrong, and that is exactly the error that surfaces at BAS time.

None of these are reasons to distrust the tool. They are reasons to point your attention where it counts. A bookkeeper who knows the AI is rock solid on recurring vendors and shaky on new ones, mixed purchases, and GST can review a month's coding in a fraction of the time it would take to check every line, and catch more than a tired human eyeballing everything would. That is not a limitation of AI. It is how you use it well.

## Building a Workflow You Can Actually Trust

Trust does not come from an accuracy statistic on a vendor's website. It comes from a process that assumes the AI will occasionally be wrong and is built to catch it before it matters. The name for this is human-in-the-loop, and it is the pattern that separates finance teams who quietly rely on AI from the ones still burned by it. Three steps.

### 1\. Let the AI draft everything

The AI codes every transaction and proposes the GST treatment, doing the volume work that used to eat the week. Crucially, it also flags what it is unsure about, so the output is not a flat list but a sorted one: confident here, uncertain there.

### 2\. Review the exceptions, not the lot

A human checks the flagged items and the known failure modes: new vendors, ambiguous and mixed transactions, one-offs, and GST edge cases. This is skilled, focused work, not data entry. It is where your judgement and your knowledge of the client actually earn their keep.

### 3\. Feed the corrections back

Every correction teaches the system. Code a new vendor once and it remembers. Set a rule for a tricky GST line and it applies next time. The pool of exceptions shrinks month over month, so the workflow gets faster and more trustworthy the longer it runs.

Worth Remembering

This is why generic, one-size-fits-all automation frustrates people while [dedicated AI](/blog/what-is-ai-agent/) earns trust. A tool trained on your firm's chart of accounts, your recurring suppliers, and your GST preferences gets the firm-specific coding right far more often than a default engine guessing from scratch. It still hands the edge cases to a human, but there are fewer of them, and the ones that remain are genuinely worth a person's time. When we ran dedicated AI over a real client's books, [it caught coding a rushed human had missed](/blog/ai-bookkeeper-catches-what-humans-miss/) and handed it back for a decision.

## The Bottom Line

Can you trust AI to categorise transactions correctly? Yes, but the trust is in the process, not the software on its own. Let the AI draft, because it is fast and genuinely good at the routine bulk. Keep a human on the exceptions, because the last few percent are where the money and the GST risk sit. Feed the corrections back, so the whole thing gets sharper every month. Do that and you get the speed of automation with the accountability of a professional sign-off, which is the combination clients are actually paying for.

The finance teams struggling with AI are usually the ones who expected blind faith to work, or who never set up a review step at all. The ones quietly thriving treat AI as a fast, tireless first pass and themselves as the judgement on top. If you want to see what a trustworthy, human-in-the-loop setup looks like for your books, that is exactly what we help firms build. [Book a free consultation](/contact) and we will show you how dedicated AI drafts the coding, flags the exceptions, and learns your firm-specific rules, while you stay firmly in the chair that signs it off.
