AI Accounting Software RFP: 12 Audit Trail Questions
12 audit trail questions to paste into an AI accounting software RFP, with a good answer and a red flag for each, a scoring sheet, and a copy-paste block covering replay, overrides, SOC 1 scope and CUECs.

Most AI accounting software RFPs ask about integrations, pricing and security, then stop. The questions that decide whether your auditor can rely on the output are usually missing, and these 12 fill that gap.
TL;DR: Add 12 audit trail questions to your AI accounting software RFP: input lineage, rule and model version, replay, uncertainty handling, overrides, reviewer evidence, access and change logs, exceptions, SOC 1 Type 2 scope, subservice organizations and CUECs, model change notice, and export and retention. Score each answer as good or red flag before any demo.
Why does an AI accounting software RFP need audit trail questions?
Because your auditor will ask them later, and it is cheaper to ask the vendor now. An AI audit trail is the record that lets someone other than the vendor see what went into an output, what logic produced it, who changed it and who approved it.
Three things raised the bar in the last two years. The PCAOB amended AS 1105 and AS 2301 to address audit procedures that use technology-assisted analysis of information in electronic form, effective for audits of fiscal years beginning on or after December 15, 2025 (PCAOB). That means calendar 2026 audits of public companies. COSO published Achieving Effective Internal Control Over Generative AI in February 2026, mapping gen AI risks to the five COSO components (COSO). And the NIST AI Risk Management Framework, released in January 2023 and voluntary, organizes AI risk work into Govern, Map, Measure and Manage.
None of these tell you which vendor to buy. They do tell you what evidence your auditor will expect to exist. If you are private and not under PCAOB standards, the logic still applies: your auditor needs to rely on the data and logic behind the numbers.
If you already read our five controls in why finance teams are still wary of giving AI control or the six evaluation questions in AI journal review needs governance, this post is the procurement version: questions written to paste into an RFP and score.
What should you ask about how the output was produced?
1. Can you show the exact source records behind any output?
Why it matters: Input lineage is the first thing an auditor traces. A match, a coded invoice or a flagged entry is only as good as the records it came from.
Good answer: Every output links to source record IDs and a snapshot or hash of the input as it was at run time, not as it is today.
Red flag: "The model had access to your bank feed and ERP." Access is not lineage.
2. Which rule version and model version produced this output, and is that stored with it?
Why it matters: Logic changes. If a match rule was edited in October, a September match must still show the September rule.
Good answer: Each output stores a rule ID, rule version, model identifier and prompt or configuration version, all queryable.
Red flag: "We always run the latest model." That makes last quarter's outputs impossible to explain.
3. Can we or our auditor rerun a past decision and get the same result?
Why it matters: Reperformance is a standard audit procedure. If the vendor cannot replay a decision with its original inputs and logic, the auditor has to test around the tool.
Good answer: A documented replay that uses the stored input snapshot and stored logic version, with a note on which steps are deterministic and which are not.
Red flag: "Results may vary slightly each run" with no way to pin inputs and logic.
Here is what a replayable record can look like. The data is illustrative, not from a real customer.
-- Illustrative: pull everything needed to reperform one AI match
SELECT r.run_id, r.run_at, r.rule_id, r.rule_version,
r.model_version, r.input_snapshot_id,
r.decision, r.reason_text,
o.override_by, o.override_reason, o.approved_by, o.approved_at
FROM match_runs r
LEFT JOIN overrides o ON o.run_id = r.run_id
WHERE r.source_txn_id = 'BANK-2026-09-000481';
If a vendor can produce an equivalent of this for any output in your sample, questions 1, 2, 3 and 5 are mostly answered in one screen.
4. What does the system do when it is uncertain?
Why it matters: The dangerous output is a confident wrong one. You want the tool to stop and route work to a person, and to say why.
Good answer: Defined thresholds, a reason attached to each low-confidence item, and a queue where those items wait for a named reviewer. For where a model should stop and hand off, see ten accounting use cases for typed decisions.
Red flag: A single confidence score with no explanation, or "the model rarely gets it wrong."
Want replayable outputs in your own close? Join the Yoraito waitlist for early access.
What should you ask about people, changes and exceptions?
5. How are overrides captured?
Why it matters: Overrides are where judgment enters, and where management override risk lives.
Good answer: The original AI output is kept, the override is a separate record with user, timestamp and required reason, and overrides can be reported by user and by period.
Red flag: The override replaces the original value, or the reason field is optional.
6. What evidence shows a named person reviewed and approved the output?
Why it matters: A review control that leaves no evidence did not happen, as far as testing is concerned.
Good answer: Reviewer identity, timestamp and what they reviewed are stored with the output, and preparer and approver cannot be the same user.
Red flag: "Users can mark items as reviewed" with no segregation of duties enforcement.
7. Who can change rules, thresholds and user access, and is every change logged?
Why it matters: Access and change management are core IT general controls. A tool that lets anyone edit match rules undermines every output downstream.
Good answer: Role-based permissions, an append-only log of configuration and access changes with before and after values, and a report you can hand to your auditor.
Red flag: Admin changes are visible only to the vendor's support team.
8. How are exceptions handled and aged?
Why it matters: Unmatched items and failed checks are where misstatements hide. An exception that silently drops out of the queue is worse than no automation.
Good answer: Every exception has an owner, a status, an age and a resolution note, and the count of open exceptions ties to period-end reports.
Red flag: Exceptions exist only as a UI list that cannot be exported or reconciled.
What should you ask about assurance and the vendor relationship?
9. Do you have a SOC 1 Type 2 report, a SOC 2 Type 2 report, or both, and what is in scope?
Why it matters: A SOC 1 report covers controls at a service organization relevant to your internal control over financial reporting; a SOC 2 report covers security, availability, processing integrity, confidentiality or privacy. The AICPA describes SOC 1 as intended for user entities and the CPAs who audit them (AICPA). If the tool produces numbers in your financial statements, your auditor will usually ask for SOC 1. Type 2 means the auditor tested whether controls operated over a period, not only whether they were designed.
Good answer: The report type, period, auditor and in-scope system are stated, and the AI features you are buying are inside the system description.
Red flag: "We are SOC 2 compliant" with no report, no period, or with the AI module out of scope.
10. Which subservice organizations do you use, and what user entity controls do you expect from us?
Why it matters: Most AI tools rely on cloud hosts and model providers. Complementary user entity controls (CUECs) are the controls the vendor's report assumes you perform; if you do not perform them, the report's conclusions may not hold for you.
Good answer: A list of subservice organizations, including any third-party model provider, whether each is carved out or included, and a written CUEC list you can map to your own controls.
Red flag: The model provider is not named, or CUECs are "standard" and not listed.
11. How and when will you notify us of model or logic changes?
Why it matters: A model swap can change outputs without any change on your side. Your change management control needs a trigger.
Good answer: Advance written notice for model or material logic changes, release notes that name what changed, and the ability to compare outputs before and after.
Red flag: "We continuously improve the model" with no notice commitment in the contract.
12. Can we export the full audit trail, and how long do you keep it?
Why it matters: Your retention obligations do not end when the contract does. Auditors may ask about a prior period after you switch tools.
Good answer: Full export of outputs, inputs references, versions, overrides, approvals and logs in an open format, retention terms written into the contract, and an exit export at termination.
Red flag: Export covers outputs only, or logs are kept for a short rolling window.
How do you score vendor answers?
Score each question 2 for a good answer, 1 for partial, 0 for a red flag or no answer. Ask for evidence (a screenshot, a sample export, a report section) rather than a yes. Use this as your scoring sheet:

- Input lineage: good, source IDs plus run-time snapshot; red flag, "it has access to your data"
- Rule and model version: good, stored with every output; red flag, "always the latest model"
- Replay: good, rerun with stored inputs and logic; red flag, results vary, nothing pinned
- Uncertainty: good, thresholds, reasons, review queue; red flag, bare score or "rarely wrong"
- Overrides: good, original kept, reason required; red flag, overwritten or optional reason
- Reviewer evidence: good, named approver, SoD enforced; red flag, anyone can mark reviewed
- Access and change logs: good, append-only, before and after; red flag, vendor-only visibility
- Exceptions: good, owner, age, ties to period end; red flag, UI list only
- SOC scope: good, SOC 1 Type 2 with AI in scope; red flag, "SOC 2 compliant," no report
- Subservice orgs and CUECs: good, named, CUECs listed; red flag, model provider not named
- Model change notice: good, advance notice in contract; red flag, "we continuously improve"
- Export and retention: good, full trail, exit export; red flag, outputs only, short window
A vendor that scores well on questions 1 to 3 and poorly on 9 and 10 may have good product design and weak assurance. The reverse, a clean SOC 2 and no replay, is the more common gap. For context on why buyers weigh this, our 2026 trust in finance AI statistics include KPMG's figure that 42% of organizations are fully assurance-ready.
What does the copy-paste RFP block look like?
Paste this into the technical or compliance section of your RFP. Ask vendors to answer each item and attach evidence.

- Q1. For any AI-generated output, provide the source record IDs and a run-time snapshot or hash of inputs. Attach a sample.
- Q2. Confirm that rule version, model identifier and configuration version are stored with each output and queryable by the customer.
- Q3. Describe how the customer or its auditor can reperform a past decision using the original inputs and logic. State which steps are non-deterministic.
- Q4. Describe confidence thresholds, how low-confidence items are routed, and whether a reason is recorded for each.
- Q5. Confirm that overrides preserve the original output and require user, timestamp and reason.
- Q6. Describe reviewer and approver evidence and how segregation of duties is enforced.
- Q7. Describe role-based access and provide a sample of the configuration and access change log.
- Q8. Describe exception ownership, aging and how open exceptions tie to period-end reporting.
- Q9. State whether you hold a SOC 1 Type 2 report, a SOC 2 Type 2 report, or both, the period, and whether the AI features are in scope.
- Q10. List subservice organizations, including model providers, with carve-out or inclusive treatment, and list all complementary user entity controls.
- Q11. State your contractual notice period and process for model or material logic changes.
- Q12. Describe export of the full audit trail, the format, retention period and exit export at termination.
Frequently asked questions
What is the difference between SOC 1 and SOC 2 for AI accounting software?
SOC 1 covers controls relevant to your internal control over financial reporting and is what your financial statement auditor usually wants. SOC 2 covers security, availability, processing integrity, confidentiality or privacy. For software that produces accounting entries or reconciliations, ask for SOC 1 Type 2 and check that the AI features are in scope.
What are complementary user entity controls?
CUECs are controls the vendor's SOC report assumes you perform, such as reviewing user access or approving outputs. If you do not perform them, the report may not support reliance for your company. Ask for the list during the RFP, not after signing.
Do private companies need to worry about PCAOB AS 1105 amendments?
The amendments apply to PCAOB audits, which cover public companies and broker-dealers. Private company audits follow AICPA standards. Still, any auditor relying on data from an AI tool needs to evaluate its reliability, so the same evidence helps either way.
Should an AI vendor guarantee identical results on rerun?
Not always, since some model steps are not deterministic. What you need is a stored record of the original inputs, logic version and output, so the decision can be explained and reperformed, plus clarity on which steps can vary.
What should you do next?
Paste the 12 questions into your RFP, score answers on evidence, and bring your auditor into the review of questions 9 and 10 before you sign. Yoraito is built so the logic behind questions 1 through 3 is SQL your team and auditor can read and rerun. To hear when it opens up, join the waitlist.