Darko Pavic - Global Retail & Fiscalization Expert

AI in Tax Compliance Needs a Verification Layer

Loading the Elevenlabs Text to Speech AudioNative Player...

New IRS guidance makes the responsibility clear: AI can assist tax professionals, but consequential output still has to be verified. The harder problem is how to make that verification scalable, traceable and usable by software and AI agents.


The IRS has put the real AI problem into writing

Much of the discussion around artificial intelligence in tax has focused on how much work the technology can automate. In June, the U.S. Internal Revenue Service pushed a more important problem to the front: the output still has to be proven correct.

On June 24, 2026, the IRS Office of Professional Responsibility issued Alert 2026-19, “Introductory Guidelines for Responsible AI Use in Federal Tax Practice.” The guidance does not create a separate AI rulebook. It applies existing professional obligations under Circular 230 to the use of AI by tax practitioners, including duties around competence, diligence, written advice, confidentiality and supervision. The practical message is clear: AI may assist professional work, but responsibility remains with the professional who relies on the result.

The part that interests me most is accuracy. The guidance, as summarized by the National Law Review and AICPA, requires practitioners to independently verify AI-generated facts, citations and calculations and to treat AI-generated text as a draft rather than a finished professional work product, which means that human oversight, professional judgment and accountability do not disappear because the first version was produced by a model.

That position is reasonable and, in a professional environment, necessary, but it also exposes the problem that I believe tax and compliance technology now has to solve.

Human review is necessary, but it is not the scalable answer

If every AI-generated tax or compliance answer has to be checked again from the beginning by a qualified expert, AI can still save time in research, drafting and summarization, but the most important bottleneck remains almost unchanged because the system may generate an answer in seconds while a professional still has to identify the correct source, confirm that it is current, determine whether it applies to the exact situation, check exceptions and verify that the conclusion follows from the rule.

I do not see this as an argument against human responsibility; I would argue the opposite, because human responsibility becomes more important as AI is used for more consequential work, but human responsibility should not mean that every professional has to manually repeat the complete legal and technical analysis behind every machine-generated result.

The real opportunity is to build a verification layer between AI and the action that follows from it, not to replace the qualified professional, the legislator, the tax authority or the court, but to make the path from source to conclusion more structured, more visible and easier to audit.

This distinction becomes important as AI moves from producing text to making or supporting decisions. An AI-generated explanation that a professional reads as background material is one thing, while an output that determines how a transaction is treated, which VAT rule applies, whether a fiscal receipt has to be created, whether a transaction may be processed offline or which reporting obligation has been triggered is something very different, and the closer AI moves to execution, the less useful it is to say only that a human should check the answer afterwards.

The problem is not access to more documents

The instinctive answer to unreliable AI is often to give the model more documents. Better retrieval, larger document collections and retrieval-augmented generation can certainly improve results, but access to the right PDF is not the same as proving that a compliance decision is correct.

NIST’s Generative AI Profile describes “confabulation” as confidently presented erroneous or false content and explains that it can arise from the way generative models produce outputs from statistical patterns, while also warning that the risk deserves particular attention when generative AI is integrated into applications involving consequential decision-making.

The OECD reaches a similar conclusion from the perspective of tax administration. Its report on AI in tax administration says that a lack of transparency, explainability and interpretability can create risks for the rule of law and taxpayers’ legal recourse, and it points to governance, oversight, data quality, validation and regular updates as necessary foundations for trustworthy AI use.

The problem is therefore wider than hallucination. Even a perfectly retrieved legal document may not tell a system whether it is the current version, whether a transition period applies, whether the rule covers the legal entity executing the transaction, whether the customer is a business or a consumer, whether an exception overrides the general rule or whether authority guidance has changed the practical interpretation without changing the underlying law.

I wrote about this boundary earlier in “Can LLMs be trusted in compliance?”, where my main concern was that fluency is not the same as legal reliability. A model can explain a rule very well and still be unsuitable as the final decision engine when the result must be repeatable, auditable and defensible.

Compliance intelligence comes before compliance automation

This is why I believe the next generation of compliance knowledge platforms has to move beyond document storage and search. The useful unit of knowledge is no longer only a document; it is a governed piece of compliance knowledge that carries its authority, jurisdiction, legal status, version, effective dates, applicability, exceptions, review status and relationship to the business process in which it will be used.

I use the term Compliance Intelligence for this broader idea. The Compliance Intelligence Library on my website describes the transition from documents to structured, trusted and implementation-oriented knowledge. The goal is not simply to make a regulation searchable, but to make the knowledge around that regulation understandable by humans, software and AI systems without losing the context that gives the rule its meaning.

In the architecture I have been exploring, vector databases and semantic search can help locate material by meaning rather than only by matching keywords, knowledge graphs can connect authorities, concepts, transactions, actors and dependencies, and metadata can identify jurisdiction, source status, publication date, effective date, review status and version. None of those components alone proves that a transaction is compliant, but together they can create a much stronger foundation than a folder of PDFs and a chatbot sitting on top of it.

For AI, this difference is fundamental because a model should not only receive a paragraph that looks relevant; it should also know whether that paragraph comes from a binding law, authority guidance, a technical specification or an internal interpretation, whether the information is current, which business facts determine applicability and where uncertainty requires human escalation instead of an automatic answer.

From an answer to a provable decision

The next step is what I call a Compliance Compiler, a research concept I have been developing around machine-readable law and regulated software. The idea is to create a controlled transformation from authoritative sources into validated, versioned and machine-readable compliance instructions, tests and evidence requirements, while the executing system may be a traditional POS, an e-commerce platform, an ERP application or, increasingly, an AI agent and the compliance rule remains external, traceable and governed.

This distinction matters because AI should not become both the actor and the authority that defines whether its own action is legal. A model can help interpret, compare, draft and orchestrate, but a regulated action needs a rule layer that the model cannot simply invent or silently change. I describe this architecture in more detail in “Compliance Compiler and Machine-Readable Law.”

The verification problem can then be expressed much more precisely. Instead of asking a professional to read the entire answer again and decide whether it sounds right, the system should be able to show which source supports the result, which version was used, which facts made the rule applicable, which exception was considered, which test was executed and what evidence was created, while an ambiguous rule or a missing approved interpretation should trigger escalation rather than a confident answer.

That approach does not remove human judgment; it gives human judgment a better place in the architecture, where experts spend their time validating interpretations, resolving ambiguity, approving rule changes and handling exceptional cases instead of repeatedly checking whether a model copied a date correctly or selected the right paragraph from a long document.

Retail makes the problem visible very quickly

Retail is a useful environment for understanding why this matters because compliance often sits directly in the transaction path. A fiscal requirement can determine the content of a receipt, the sequence of a transaction, a signature, a reporting message, the treatment of a return, the behavior during an outage or the evidence that has to be stored, which means that the output is not an abstract legal summary but system behavior.

The same issue will become more visible as AI agents enter commerce. An agent may search for a product, negotiate, order, return, refund or initiate another commercial action, but the legal obligations attached to that action do not disappear because the interface is conversational or autonomous; the agent still operates in a jurisdiction, for a legal entity, through a channel, with a customer type, payment method and transaction context that can change which rule applies.

For that reason, I do not think the long-term solution is to teach a general AI model every fiscal rule in the world and hope it remembers them correctly. The more robust architecture is to let general AI remain general while connecting it to a maintained compliance intelligence layer that can provide validated knowledge and, eventually, a separate mechanism that can test the resulting action against the applicable rules.

This is where the IRS guidance becomes much more interesting than a professional-ethics note for U.S. tax practitioners, because it makes the verification requirement explicit. The IRS position is that an AI-generated result is not self-validating and that the professional remains responsible for accuracy, while AICPA’s summary of the guidance reinforces the same point through professional judgment, due diligence, confidentiality safeguards and accountability. [1][4] The technology industry now has to answer the operational part of that problem.

The research area is becoming clearer

We are currently working on this problem and defining the research areas that have to be solved before AI can be trusted more deeply in compliance-critical processes. The work touches legal informatics, knowledge representation, provenance, semantic technologies, formal and executable rules, software testing, AI evaluation, synthetic data and governance, while the difficult part is not any single technology but the controlled transformation between legal authority and operational behavior.

One research area is source authority and provenance, because a system has to distinguish between law, authority guidance, technical specifications, professional interpretation and internal implementation advice. Another is applicability, because the same rule can produce a different result depending on jurisdiction, entity, transaction type, customer type, product, payment method, channel and time. A third area is verification itself, including how a legal obligation can become a testable rule, how uncertainty is represented, how rule changes trigger regression testing and how the resulting evidence can later be audited.

Synthetic data is another part of the problem that I find particularly interesting. If verified scenarios can be derived from validated rules, including positive cases, negative cases, exceptions, boundary conditions and failure situations, those scenarios could become a much stronger basis for training and evaluating AI systems than unverified examples collected from the open internet. The research question is not simply how to create more synthetic data, but how to create synthetic compliance data whose expected result can be traced back to an authoritative rule and an approved interpretation.

We are currently looking for cooperation around these questions with universities, researchers, AI companies, tax and legal specialists, retailers and software providers. The subject is too interdisciplinary to be solved by one profession alone because the legal interpretation has to survive the transformation into data structures, software behavior, tests and evidence without losing its authority along the way.

The missing layer is verification

The IRS guidance is useful because it separates enthusiasm for AI from responsibility for the outcome. It does not tell tax professionals to avoid AI; it says that the professional obligations remain when AI is used and that the output has to be verified before it becomes professional work.

I think this will become one of the defining problems of AI in tax and compliance. Models will continue to improve, retrieval will become better, context windows will grow and agents will perform more work independently, but none of those developments removes the need to prove that a regulated action is based on the correct rule, in the correct version, for the correct situation.

The next important step is therefore not only better AI, but better compliance knowledge and a reliable mechanism for turning that knowledge into constraints, tests and evidence that both humans and machines can use, which is the layer I believe we now have to build.

Sources and further reading

1. Internal Revenue Service, Office of Professional Responsibility, “Introductory Guidelines for Responsible AI Use in Federal Tax Practice,” Alert 2026-19, June 24, 2026. https://content.govdelivery.com/accounts/USIRS/bulletins/41d6e70

2. Internal Revenue Service, “Office of Professional Responsibility and Circular 230.” https://www.irs.gov/tax-professionals/office-of-professional-responsibility-and-circular-230

3. Joseph J. Lazzarotti, Melissa Ostrower and Jackson Lewis P.C., “IRS Office of Professional Responsibility (OPR) Issues AI Guidance: Tax Professionals Also Face AI Ethics and Compliance Obligations,” National Law Review, August 4, 2026. https://natlawreview.com/article/irs-office-professional-responsibility-opr-issues-ai-guidance-tax-professionals

4. AICPA & CIMA, “AICPA insights: Introductory guidelines for responsible AI use in federal tax practice,” July 2, 2026. https://www.aicpa.com/resources/article/aicpa-insights-introductory-guidelines-for-responsible-ai-use-in-federal-tax-practice

5. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile,” NIST AI 600-1, July 2024. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

6. OECD, “AI in tax administration: Governing with Artificial Intelligence,” 2025. https://www.oecd.org/en/publications/2025/06/governing-with-artificial-intelligence_398fa287/full-report/ai-in-tax-administration_30724e43.html

Related reading on darkopavic.xyz

Compliance Intelligence Library

Can LLMs be trusted in compliance?

Compliance Compiler and Machine-Readable Law

Darko Pavic

Darko Pavic is a retail technology and fiscalization expert with more than 28 years of experience in international POS systems, retail compliance and software architecture. His current work focuses on fiscalization, e-invoicing, compliance intelligence, machine-readable regulation and the responsible use of AI in compliance-critical systems.

https://darkopavic.xyz