FIELD GUIDE / FINANCE AI

Questions to ask an AI vendor.

A controller's guide to evaluating inputs, repeatability, uncertainty, human review, audit trails, and data handling. Ask these before a pilot, not after.

FINANCE PRACTICEField notesYORAITO / RESOURCE DESK
FOR CONTROLLERS & EVALUATION TEAMS12 QUESTIONS / 6 GROUPS
Printable working copy

Judge the control, not the demo.

Demos show the happy path. These questions target the parts of a finance workflow that decide whether the output can be relied on: where the data came from, what happens when the model is unsure, and what a reviewer can see afterwards.

QUESTIONS

Inputs and data handling

Q

What data does the system read, and what does it retain?

Identify the exact sources: ledger extracts, bank feeds, invoices, contracts, and spreadsheets. Confirm whether data is used to train any model, how long it is retained, and whether it can be deleted.

STRONG ANSWER

Names the specific systems it reads, states retention in writing, and can delete workspace data on request.

RED FLAG

Answers only in general terms, or treats 'we don't train on your data' as the complete answer without naming subprocessors.

Q

How are credentials and permissions scoped?

Ask which actions are read-only and which can write back to a system of record. A vendor that can post a journal is operating inside your control environment.

STRONG ANSWER

Separates read and write scopes, supports scoped credentials, and logs every write with the requesting user.

RED FLAG

Uses a single shared administrator credential, or offers posting without a named approval step.

QUESTIONS

Determinism and repeatability

Q

Is the same input going to produce the same output?

Finance work needs to be re-performable. Ask which steps are deterministic code and which steps send the data to a model, and how the second run of an unchanged input is verified.

STRONG ANSWER

Uses deterministic transformations for structured work and reserves model calls for genuinely ambiguous cases, with recorded outputs.

RED FLAG

Describes the whole pipeline as 'AI-powered' without separating logic from inference, so results cannot be reproduced.

Q

How do you handle a change in model or prompt behaviour?

Models, prompts, and tool versions change. Ask how a vendor detects that output quality shifted and who is notified.

STRONG ANSWER

Versions the workflow, pins the model where feasible, and monitors output against expected checks after any change.

RED FLAG

Has no answer for detecting a regression, or treats a model upgrade as a routine deployment.

QUESTIONS

Exceptions and uncertainty

Q

What happens when the system is not confident?

This is the single most important question. Ask what the threshold is, who is notified, and whether the item can proceed without human input.

STRONG ANSWER

Routes low-confidence items to a named owner with the evidence attached, and blocks irreversible steps until reviewed.

RED FLAG

Always produces an answer, presents a guess as a result, or quietly excludes difficult items from the output.

Q

How are edge cases discovered after deployment?

Ask how the vendor finds the inputs that break the workflow, and whether customers can see the exception population rather than only the successes.

STRONG ANSWER

Surfaces an exception queue with age and owner, and shares patterns across comparable customers.

RED FLAG

Reports only completion rates, or treats errors as individual user problems.

QUESTIONS

Review and approval

Q

Can a reviewer check the work without redoing it?

A reviewer needs the source, the logic, and the open questions. Ask to see a completed workflow the way an auditor would encounter it.

STRONG ANSWER

Shows a per-item record with source, reasoning or rule applied, exceptions raised, and the approving user.

RED FLAG

Provides only a final output file, or a dashboard with no link back to the underlying items.

Q

Who is accountable for a wrong result?

Establish in writing that your team retains sign-off responsibility and understand the vendor's contractual position on errors and corrections.

STRONG ANSWER

Documents responsibilities clearly and has a process for correcting a defect that reached a customer's ledger.

RED FLAG

Offers no remediation path, or implies the customer cannot verify results independently.

QUESTIONS

Audit trail and evidence

Q

What is retained, and can it be exported?

Retention that cannot be exported is not evidence. Ask for the audit record format, the retention period, and whether it can leave the platform in a readable form.

STRONG ANSWER

Exports a complete run record, including inputs, intermediate steps, exceptions, and approvals.

RED FLAG

Keeps the trail only inside the product, or retains it for a period shorter than your audit requirement.

Q

Can a third party reconstruct the result?

Ask whether an external reviewer could follow the trail from source to final number without a demonstration from the vendor.

STRONG ANSWER

The trail stands on its own, and the vendor supports a formal review process.

RED FLAG

Requires a live walkthrough, or relies on the vendor's staff to explain the output.

QUESTIONS

Security, continuity, and change

Q

Which subprocessors handle our data, and how are we notified of changes?

Ask for the named list of subprocessors, the data each one holds, and the notice period before a new one is added. This is also the question to ask about data residency.

STRONG ANSWER

Maintains a published subprocessor list, gives contractual notice before changes, and names the region data is held in.

RED FLAG

Cannot name the subprocessors, treats 'we use reputable providers' as sufficient, or gives no notice of a change.

Q

What happens to our data and audit trail if we terminate?

Establish what you can take with you and what is deleted. The audit trail matters as much as the working data, particularly if you are still defending a close period.

STRONG ANSWER

Provides a documented export of both the work product and the run history, with a defined deletion schedule and written confirmation.

RED FLAG

Offers no export path, retains data indefinitely, or cannot separate your records from other customers'.

PILOT

A pilot shape that produces real evidence

Most evaluations fail because they measure volume instead of risk. Structure the pilot around the questions you actually need answered.

A pilot shape that produces real evidence
Pilot elementWhat to defineWhat it tells you
ScopeOne recurring workflow, one entity, one periodWhether the tool fits a real process rather than a curated demo input
PopulationEvery item in scope, including the hard onesThe true exception rate, which is what review cost depends on
ComparisonThe current manual outcome, prepared independentlyWhether time and rework actually fall, and where they move instead
ReviewA controller reviews output before anything is postedWhether the output is checkable in the time you actually have
EvidenceThe exported run record, retained after the pilotWhether the trail would satisfy a review without vendor help
Exit criteriaWritten before the pilot startsA defensible decision instead of a subjective impression
THE POINT OF THE EXERCISE

You are evaluating a control, not a model.

The vendor's accuracy on a sample is not the decision. The decision is whether you can verify the work, catch the exceptions, and stand behind the number. If those three things hold, the tooling matters less than the current process.

EARLY ACCESS

Evaluating a finance AI tool?

We work through evaluation criteria with finance teams. Bring the vendor's answers and we will tell you what is missing.

Start a conversation