Questions to ask an AI vendor.
A controller's guide to evaluating inputs, repeatability, uncertainty, human review, audit trails, and data handling. Ask these before a pilot, not after.
Judge the control, not the demo.
Demos show the happy path. These questions target the parts of a finance workflow that decide whether the output can be relied on: where the data came from, what happens when the model is unsure, and what a reviewer can see afterwards.
Inputs and data handling
What data does the system read, and what does it retain?
Identify the exact sources: ledger extracts, bank feeds, invoices, contracts, and spreadsheets. Confirm whether data is used to train any model, how long it is retained, and whether it can be deleted.
Names the specific systems it reads, states retention in writing, and can delete workspace data on request.
Answers only in general terms, or treats 'we don't train on your data' as the complete answer without naming subprocessors.
How are credentials and permissions scoped?
Ask which actions are read-only and which can write back to a system of record. A vendor that can post a journal is operating inside your control environment.
Separates read and write scopes, supports scoped credentials, and logs every write with the requesting user.
Uses a single shared administrator credential, or offers posting without a named approval step.
Determinism and repeatability
Is the same input going to produce the same output?
Finance work needs to be re-performable. Ask which steps are deterministic code and which steps send the data to a model, and how the second run of an unchanged input is verified.
Uses deterministic transformations for structured work and reserves model calls for genuinely ambiguous cases, with recorded outputs.
Describes the whole pipeline as 'AI-powered' without separating logic from inference, so results cannot be reproduced.
How do you handle a change in model or prompt behaviour?
Models, prompts, and tool versions change. Ask how a vendor detects that output quality shifted and who is notified.
Versions the workflow, pins the model where feasible, and monitors output against expected checks after any change.
Has no answer for detecting a regression, or treats a model upgrade as a routine deployment.
Exceptions and uncertainty
What happens when the system is not confident?
This is the single most important question. Ask what the threshold is, who is notified, and whether the item can proceed without human input.
Routes low-confidence items to a named owner with the evidence attached, and blocks irreversible steps until reviewed.
Always produces an answer, presents a guess as a result, or quietly excludes difficult items from the output.
How are edge cases discovered after deployment?
Ask how the vendor finds the inputs that break the workflow, and whether customers can see the exception population rather than only the successes.
Surfaces an exception queue with age and owner, and shares patterns across comparable customers.
Reports only completion rates, or treats errors as individual user problems.
Review and approval
Can a reviewer check the work without redoing it?
A reviewer needs the source, the logic, and the open questions. Ask to see a completed workflow the way an auditor would encounter it.
Shows a per-item record with source, reasoning or rule applied, exceptions raised, and the approving user.
Provides only a final output file, or a dashboard with no link back to the underlying items.
Who is accountable for a wrong result?
Establish in writing that your team retains sign-off responsibility and understand the vendor's contractual position on errors and corrections.
Documents responsibilities clearly and has a process for correcting a defect that reached a customer's ledger.
Offers no remediation path, or implies the customer cannot verify results independently.
Audit trail and evidence
What is retained, and can it be exported?
Retention that cannot be exported is not evidence. Ask for the audit record format, the retention period, and whether it can leave the platform in a readable form.
Exports a complete run record, including inputs, intermediate steps, exceptions, and approvals.
Keeps the trail only inside the product, or retains it for a period shorter than your audit requirement.
Can a third party reconstruct the result?
Ask whether an external reviewer could follow the trail from source to final number without a demonstration from the vendor.
The trail stands on its own, and the vendor supports a formal review process.
Requires a live walkthrough, or relies on the vendor's staff to explain the output.
Security, continuity, and change
Which subprocessors handle our data, and how are we notified of changes?
Ask for the named list of subprocessors, the data each one holds, and the notice period before a new one is added. This is also the question to ask about data residency.
Maintains a published subprocessor list, gives contractual notice before changes, and names the region data is held in.
Cannot name the subprocessors, treats 'we use reputable providers' as sufficient, or gives no notice of a change.
What happens to our data and audit trail if we terminate?
Establish what you can take with you and what is deleted. The audit trail matters as much as the working data, particularly if you are still defending a close period.
Provides a documented export of both the work product and the run history, with a defined deletion schedule and written confirmation.
Offers no export path, retains data indefinitely, or cannot separate your records from other customers'.
A pilot shape that produces real evidence
Most evaluations fail because they measure volume instead of risk. Structure the pilot around the questions you actually need answered.
| Pilot element | What to define | What it tells you |
|---|---|---|
| Scope | One recurring workflow, one entity, one period | Whether the tool fits a real process rather than a curated demo input |
| Population | Every item in scope, including the hard ones | The true exception rate, which is what review cost depends on |
| Comparison | The current manual outcome, prepared independently | Whether time and rework actually fall, and where they move instead |
| Review | A controller reviews output before anything is posted | Whether the output is checkable in the time you actually have |
| Evidence | The exported run record, retained after the pilot | Whether the trail would satisfy a review without vendor help |
| Exit criteria | Written before the pilot starts | A defensible decision instead of a subjective impression |
You are evaluating a control, not a model.
The vendor's accuracy on a sample is not the decision. The decision is whether you can verify the work, catch the exceptions, and stand behind the number. If those three things hold, the tooling matters less than the current process.
Evaluating a finance AI tool?
We work through evaluation criteria with finance teams. Bring the vendor's answers and we will tell you what is missing.