CVOCA

AI Model Risk and Explainability in Financial Statement Preparation

Rushabh Gala September 2, 2026 Auditing & Ethics ⏱️ 8 min read

What CAs must document when AI tools assist in judgment calls

AI tools have moved well past the experimentation phase in most mid-to-large accounting practices. They now assist in areas that were once the exclusive domain of professional judgment: classifying expenditure as capital or revenue, estimating expected credit loss allowances, flagging revenue-recognition anomalies, drafting disclosure notes, and increasingly, forming a first view on materiality and risk. The convenience is real. So is the exposure. When an AI tool contributes to a judgment call embedded in a set of financial statements, the question a CA must be able to answer — to an audit committee, a peer reviewer, a regulator, or a tribunal — is not “did the AI get it right,” but “can you show how the conclusion was reached, and that you exercised independent professional judgment over it.” This article sets out what that documentation obligation looks like in practice.

1. Why This Is Now a Live Compliance Issue, Not a Future One

Three developments have brought AI model risk from a theoretical concern to a documentation requirement with teeth. First, disclosure expectations around AI use in financial reporting and audit processes have tightened materially through 2026, with regulators increasingly expecting firms to be able to produce a record of where AI tools were used, what they influenced, and how outputs were validated — not merely a general statement that AI-assisted tools were deployed. Second, global model-risk-management guidance — most notably the revised supervisory framework that replaced the decade-old US model risk guidance in April 2026 — has pushed governance expectations for AI models well beyond the traditional “black box, but it worked in testing” standard, and these expectations are filtering into Indian practice through multinational audit networks and increasingly sophisticated audit committees. Third, ICAI’s own direction — through CA GPT’s expanding capabilities and its published guidance on AI governance within the profession — has been explicit that AI-assisted findings must be validated, documented, and subject to “effective challenge” by a qualified professional before they inform a reported figure.

Put simply: a peer reviewer or NFRA inspection team in 2026 is far more likely to ask “show me your AI validation workpaper” than they were even two years ago.

2. What Counts as an “AI-Assisted Judgment Call”

Before documentation can be designed, the scope of what needs documenting has to be defined clearly, because not every use of AI in a finance function carries the same risk. It is useful to distinguish three tiers.

Tier 1 — Mechanical/low-risk automation

Bank reconciliation matching, OCR-based data extraction from invoices, and standard journal-entry population. These involve minimal judgment and are typically governed adequately by existing IT-general-control documentation, though even here, an audit trail of what was auto-populated versus manually entered is good practice.

Tier 2 — Analytical support to judgment

Anomaly detection flagging unusual journal entries for review, AI-assisted variance analysis, or a model suggesting an expected credit loss range based on historical default patterns. Here the AI narrows a range or highlights an exception; a human still makes the final call, but the AI has materially shaped the starting point of that call.

Tier 3 — AI as a direct input into an accounting estimate or classification

A machine-learning model directly generating a provision estimate, a fair-value input, an impairment trigger assessment, or drafting disclosure language substantively adopted with limited rewriting. This tier carries the highest documentation burden, since the AI output is functionally indistinguishable, in the financial statements, from a judgment the CA would otherwise have made unaided.

The documentation standard set out below should be applied in full to Tier 3, applied proportionately to Tier 2, and largely satisfied by existing IT controls for Tier 1.

3. The Core Documentation Package

For any Tier 2 or Tier 3 use of AI, the working papers should be able to answer five questions, each backed by a specific artefact.

3.1 What tool was used, and what is it designed to do?

Document the tool name, version, vendor (or in-house build), and — critically — the tool’s stated purpose and known limitations as described by its own documentation or model card. A generic large language model used to draft a disclosure note carries very different risk characteristics from a purpose-built ECL model trained on the entity’s own historical data, and the workpaper should make that distinction explicit rather than treating “AI” as a single undifferentiated category.

3.2 What data went in, and was it appropriate?

Record the data set or inputs provided to the model — source, period, completeness, and known quality issues — along with an assessment of whether that data remains representative of current facts and circumstances. A credit-loss model trained predominantly on pre-pandemic default data needs an explicit note on why that training data remains an appropriate basis for a 2026 estimate. This is a data-lineage requirement, and its absence is among the most common gaps found in early AI-audit reviews.

3.3 What did the model output, and how confident should we be in it?

Capture the raw output — the flagged exceptions, the suggested range, the drafted language — before any human edits are applied, so that the “before AI” and “after human review” positions can both be evidenced. Where the tool provides a confidence score, probability range, or similar output, that should be recorded too, since a materially wide confidence interval is itself relevant information for whether the AI output is fit to anchor a final judgment.

3.4 How was the output explained and challenged?

This is the explainability requirement proper, and it is the piece most CA workpapers currently lack. For Tier 3 uses, documentation should show what drove the AI’s output, in terms a non-technical reviewer or regulator could follow. This does not require understanding the underlying mathematics of a gradient-boosted model or neural network; it requires the tool (or a validation layer around it) to produce an intelligible rationale — of the kind explainability techniques such as SHAP or LIME provide by decomposing a prediction into each input variable’s contribution, or a well-designed model card provides by describing intended use, training data, and known failure modes. Where a tool genuinely cannot produce such a rationale — a pure “black box” — that fact itself must be documented, along with the compensating human procedures performed to independently corroborate the output.

3.5 What did the qualified professional actually decide, and why?

The single most important line in the workpaper records the human conclusion — separately and explicitly from the AI’s suggested output — along with the professional’s reasoning for accepting, adjusting, or rejecting it. Regulators are not looking for evidence that the AI was right; they are looking for evidence that a qualified individual exercised genuine, independent professional scepticism over the output rather than rubber-stamping it. A workpaper showing the AI’s suggested figure and the final reported figure as identical, with no visible evidence of independent testing, is weaker evidence than one showing the professional queried and only partially relied on the AI output — even where the final numbers converge.

4. Governance Infrastructure Firms Should Have in Place

Beyond engagement-level documentation, firms adopting AI tools for judgment-relevant work should be building three pieces of standing infrastructure, rather than reconstructing them under audit or peer-review pressure after the fact.

•  An AI tool inventory (sometimes called a Controls Library), listing every AI tool in use across the practice, its risk tier, the engagements it touches, and its last validation date — so that a peer reviewer’s first question (“which of your engagements used AI, and for what”) has a ready answer rather than requiring a scramble across files.

•  A named responsible individual — often described as a Model Steward — accountable for validating new AI tools before firm-wide rollout, reviewing periodic performance against expectations, and maintaining the change log when a tool’s version or underlying model is updated, since a validated model can silently become an unvalidated one after a vendor update.

•  A standard workpaper template for AI-assisted judgment calls, covering the five questions in Section 3 above, so that documentation quality does not depend on how AI-literate the individual engagement team happens to be.

5. A Practical Illustration

Consider a mid-sized NBFC client where an AI-assisted model generates a suggested expected credit loss provision by segment. Adequate documentation would include: the model’s name and version; the historical loan performance data used to train it, including its date range and a note on whether recent macro conditions are adequately reflected; the model’s raw suggested provision by segment, with any confidence range provided; a breakdown showing which variables (days-past-due, sector concentration, collateral coverage) most influenced the output; the engagement team’s independent back-testing of the suggested provision against a manually assessed sample; and a final memo recording where the reported provision matched the AI suggestion, where it was adjusted, and why. That memo — not the AI’s output — is what should be signed off by the engagement partner.

Conclusion

AI tools are, on balance, making judgment calls in financial statement preparation more consistent and better evidenced than manual processes often were — provided the documentation trail is built deliberately rather than assumed. The professional standard has not changed: a CA remains accountable for the judgment embedded in the numbers they sign off on. What has changed is the evidence now expected to demonstrate that accountability was real, not nominal. Firms that build AI documentation into their standard workpaper templates now — data lineage, output capture, explainability, and a clearly evidenced human decision — will find peer review, NFRA scrutiny, and audit committee questioning considerably less stressful than firms that treat this as paperwork to be assembled only when someone finally asks for it.

[The author can be reached at rushabh@dsaca.co.in]

Share on:
Scroll to Top