Back to Blog
Tax & Compliance 10 min read
By Dr. Ash Khalilian ·

AI Audit Trails, and How to Prove What Your AI Actually Did

Forums have moved past asking whether AI is accurate. The question now is whether you can evidence it. If the ATO reviews this client in two years, can you show who or what made this entry, and on what basis?

An Australian accountant at a desk tracing a single ledger entry back through a stack of source documents and a screen of system logs, warm office lighting, no legible text on screen
A record is not what you remember about the entry. It is what you can still produce two years later.

Short Answer

An adequate AI audit trail records, at the time of the action, the input document and its identifier, the system and version that acted, the instruction or rule that fired, the source records read, the output produced, the escalation state, and the named person who approved it and when. Most tools keep only the last of those, which is why most firms cannot reconstruct an entry from three months ago.

Last reviewed: September 2026

Key takeaways

  • This is a records question, not an AI question. Australian law does not care whether a human or a system made the entry, only whether you can substantiate it.
  • The ATO already requires records showing "how (basis or method)" any estimate, determination or calculation was made.
  • Taxation Ruling TR 2018/2 has required system documentation covering structure, programs, inputs and outputs since 2018.
  • Section 30 of the Code Determination 2024 requires records kept at least 5 years, showing nature, scope, outcome, basis and method.
  • Most AI tools retain the final output and a generic "modified by API user" stamp, which answers almost nothing a reviewer will ask.
  • The log must be a by-product of the action: once the action ends, the system no longer knows which records it read.

The governance argument in Australian accounting forums changed shape this year. Through 2025 the fight was about accuracy: does the model code correctly, does it invent a figure, can you trust it near a BAS. Firms actually running agents have moved to a harder question, and it decides whether a practice can use dedicated AI at all. Not "is it accurate", but "can you evidence it". Put plainly: if the ATO reviews this client in two years, can you show who or what made this entry, and on what basis?

Why is this a record-keeping question, not an AI question?

Because Australian record-keeping obligations are written about the transaction, not the tool. Nothing in the law asks what produced an entry; it asks whether you can substantiate it and produce the evidence on request. AI record keeping is not a future problem awaiting future rules: the obligation already applies to whatever you deployed last quarter.

Start with the ATO. Its overview of record-keeping rules for business, last updated 18 June 2026, requires records of every transaction relating to tax, superannuation and registration affairs, including "any documents containing details of any election, choice, estimate, determination or calculation you make for your business's tax and super affairs, including how (basis or method) the estimate, determination or calculation was made". The basis and method are part of the record. If a model made the determination, they live inside its inputs and instructions, nowhere else.

The same page carries a rule almost no AI deployment has considered: "You need to be able to reconstruct your original data if your record-keeping system changes over time." Swap models mid-year and that is a design requirement.

What does the ATO require for electronic records?

Documentation of the system itself, required since 2018. Taxation Ruling TR 2018/2, Income tax: record keeping and access, electronic records, date of effect 14 February 2018, is a public ruling whose ruling section is legally binding. Electronic records must not be altered, must stay reconstructable if the system changes, must generally be kept five years, remain retrievable, and be in English or easily convertible.

Inside the retrieval requirement sits the most useful sentence in Australian AI governance, written before anyone worried about agents in the ledger: "Keep system documentation that explains the basic aspects of the system, including structure, programs, inputs and outputs so we can ascertain that the system is doing what it is claimed to do if required." Applied to an AI that is a specification: which system, which version, what went in, what came out, evidence it does what you claim. A vendor's marketing page is not that document.

What do TPB obligations add for registered agents?

A second, separate record: the record of the service, not just the transaction. Section 30 of the Tax Agent Services (Code of Professional Conduct) Determination 2024 requires registered tax and BAS agents to keep records that correctly record the services provided to each client, including former clients. Detail sits in TPB(GS) 52/2024 Obligation to keep proper client records of tax agent services provided, issued 23 December 2024 and last modified 30 April 2026. Records must "be retained for at least 5 years after the service has been provided", "show the nature, scope and outcome of the tax agent service provided", and "reference information reasonably considered in the provision of the tax agent service".

For complex matters it requires the reasoning behind advice, "including the basis on which, and the method by which, any calculations, determinations, or estimates used, have been made". Its minimum details include who the service was provided by, how, when, and the "date that the record was made and the date of any modifications to the record". A who, what, how, when and version-history specification, and it never mentions AI.

The AI layer sits on top. TPB(GS) 55/2026 The use of Artificial Intelligence and the Code of Professional Conduct, issued 22 July 2026, says practitioners should verify and review AI generated content throughout each step of the workflow and establish processes to understand and contest AI outputs, then adds five words: "Each of these steps should be documented." The TPB ties that to sections 30 and 40. Disclosure to clients is separate, covered in the breakdown of TPB(GS) 55/2026 for tax and BAS agents.

Does the record have to say that a machine did it?

Not in those terms. As at September 2026 no Australian rule requires a ledger line to carry an "AI generated" tag, and TPB(GS) 52/2024 describes who a service was provided by as the practitioner, employees under supervision and control, contractors, or outsourced entities. An AI tool is not a separate provider: the registered practitioner is, and accountability stays there.

What bites is basis and method. When a human reaches a figure, the evidence is a working paper and a file note; when a model reaches it, the evidence is the inputs, the instruction that fired, the records consulted and the review it passed. Keep the trail, or you cannot answer the question. For AI in an audit practice that bar is already written down.

What does an adequate AI audit trail contain?

Eight items, captured per action, not per session. This is the list Agentive works through when scoping an Australian deployment.

  1. The input documents and their identifiers. A stable reference to the exact file version consumed, such as a content hash, proving the document read is the one you hold.
  2. The system and model version. Not "our AI", but the named system and version in effect at that moment.
  3. The instruction or rule that fired. The prompt, policy or rule as it stood that day, not as it reads after you tuned it.
  4. The source records read. Which accounts, prior transactions, contract clause or bank line. This is the basis element, and the item most often missing.
  5. The output as produced. The draft entry, journal or note wording, before anyone edited it.
  6. The confidence or escalation state. Auto-applied, flagged or escalated, against which threshold. This proves the review policy was operating, not just written.
  7. The human who approved it, and when. A named identity and timestamp, separate from creation, because one event cannot evidence both production and independent review.
  8. Any subsequent edit. What changed, who changed it, when and why.

Content hashing and confidence thresholds are not legal requirements in their own right; they are engineering choices that make the legal requirements provable. Keep this log separate from your system register. Guardrail 9 of the Australian Government's Voluntary AI Safety Standard, updated 2 December 2025, wants an AI inventory and system documentation. That describes the system; section 30 and the ATO rules demand records of the service.

What will a reviewer ask, and what must your logs contain?

A review arrives not as a governance framework but as blunt questions about one transaction, years after anyone remembers it.

What a reviewer will ask What your logs must contain Source of the obligation
"Who made this entry?" The accountable practitioner, plus the system, version and account behind the draft TPB(GS) 52/2024, who the service was provided by
"What did it look at?" Identifiers for every source document and ledger record read Section 30, reference information reasonably considered
"On what basis was this calculated?" The instruction or rule that fired, and the inputs it consumed ATO basis or method rule; section 30
"Did a person check it?" An approval event with a named identity and timestamp, distinct from creation TPB(GS) 55/2026; ASA 230 for audit files
"Has it changed since?" An immutable creation record plus a dated modification history TPB(GS) 52/2024; TR 2018/2 anti-alteration
"Now the other 400 entries." A queryable log, indexed and extractable to CSV ATO digital storage; TR 2018/2 retrieval
"Does the system still do what you claim?" System documentation: structure, inputs, outputs, version history TR 2018/2 system documentation

An eighth question, where the data was processed and who consented, is a Code item 6 and Privacy Act matter before a records one: see the discussion of putting client data into AI and security and data governance for an Australian AI deployment.

What do most AI tools actually retain?

The final output, and almost nothing else, in three recognisable shapes. The output-only log leaves the ledger holding a coded transaction and a note that an integration created it, with no trace of the document, the instruction or the records read. The chat-history log holds the reasoning, but in a pane not attached to the transaction and sometimes kept for under five years. The integration-layer log records that an API token acted at 11:42pm and nothing about what the token was thinking.

Worse, because it feels like compliance, is reconstruction after the fact: testimony about evidence rather than evidence, and where an invented figure becomes unfindable, as in the post on hallucinations in financial reports.

How do I test my own audit trail this week?

Run the reconstruction test on one entry. It takes about half an hour, and it is the only honest measure of whether your firm has an audit trail or a habit of trusting one.

  1. Pick one AI-assisted entry from roughly three months ago. Not last week, and not a showcase: a middling transaction on a middling client, ideally with a GST consequence.
  2. Bar the person who made it from helping. Auditing Standard ASA 230 Audit Documentation, as amended to April 2022, sets the bar: documentation "sufficient to enable an experienced auditor, having no previous connection with the audit, to understand" the work.
  3. Produce the source document and prove it is unchanged. The exact file version the system consumed, not a later copy.
  4. Name the system and version that acted. The identifier in effect that date. If your vendor has since updated it silently, note that.
  5. Retrieve the instruction that fired and the records it read. The prompt or rule as it stood, and the specific ledger lines, accounts or clauses consulted.
  6. Show the output before human editing, then the approval. ASA 230 requires a record of who performed the work and when, and who reviewed it, the date and the extent.
  7. List every change since, with authorship and date. Then repeat on a second entry for another client, because one success can be luck.

Most Australian practices will fail at step three or step five, and that deserves saying plainly. Failing is not misconduct: it means AI entered the firm as a productivity tool and never its records architecture, the gap the AI policy template for Australian accounting firms closes.

How does Agentive log what a dedicated AI does?

By writing the log at the time of the action, as a by-product of the action itself, not as a report generated afterwards. Agentive records every action a dedicated AI takes, with the source record it read and the approval state, because a practitioner cannot supervise what they cannot reconstruct. A log written during the action is a record; one assembled later is an account of one.

The engineering reason is less obvious, and it is the observation from building this for Australian finance teams that should change how practices evaluate vendors. Once an action has finished, the system genuinely no longer knows which of the nine hundred transactions it examined; that knowledge exists only in the moment of reading. Any system offering to produce an audit trail on request, rather than emitting one as it works, is rebuilding a plausible narrative from what is left in memory. Hence the question Agentive puts to any vendor: show me the log for an action you took last month, not one you can generate now.

Agentive runs single-tenant on AWS Sydney, all inference inside Australia, data never leaving Australian borders, and client data never used to train a model, aligned to APRA CPS 234, ASIC RG 255 and TPB obligations including TPB(GS) 55/2026. That makes "where was this processed" a one-sentence answer rather than a subprocessor diagram. Ledger work sits on the bookkeeping page, compliance on the tax and compliance page, and the speed claim gets tested in the piece asking whether AI can really do an audit in seconds.

What a complete log does not do

A log proves what happened. It does not decide whether what happened was correct, and treating an audit trail as a substitute for supervision is a new way to fail an old obligation. TPB(GS) 55/2026 is direct: practitioners must exercise their own professional judgement and not rely on AI output as a substitute for their own analysis of a client's circumstances.

Retention also has a ceiling. Where the Privacy Act 1988 applies, the Australian Privacy Principles may require destruction or de-identification of client information, so decide deliberately what the log holds: it inherits the privacy profile of whatever it captures, which bites first in the bookkeeping use case.

Run the test before you widen the deployment

Of the three things the forums say to track before an agent touches live accounts, hallucination rates, escalation speed and audit trails, only one is unrecoverable. Accuracy improves next quarter; a threshold tightens tomorrow. A record never written cannot be created later, and that asymmetry should set your order of work.

So pick one entry from three months ago and rebuild it. If you can, widen the deployment. If you cannot, you have found the real prerequisite, and it is not a better model but an AI Operation Engine that writes the record while it works, so "who or what made this entry, and on what basis" is a query, not an archaeology project.

General information for Australian practices, not professional, legal or tax advice. Regulatory positions are stated as at 17 September 2026 and should be verified against the primary sources linked above. Retention periods quoted are minimums.

Bring Us One Entry and We Will Try to Reconstruct It

Agentive builds dedicated AI for Australian accounting and bookkeeping practices, single-tenant on AWS Sydney, all inference inside Australia, client data never used to train a model, and a per-action log written at the time of the action. Pick one AI-assisted entry from three months ago, bring it to the call, and we will walk the reconstruction test with you on your own records.