← Back to journal
AI and Accounting

AI Bank Reconciliation: What Every Match Should Explain

A bare confidence score tells an auditor nothing. Here is what an explained AI bank reconciliation match should show: rule, version, approver, tier with a reason, match type and an exception trail.

AI bank reconciliation cover showing an example match record with the rule, approver, match type, tolerance, tier and reason, and the line: The model proposes. Only approved rules run.

"Matched, 97% confidence" tells a controller nothing, and an auditor even less. In AI bank reconciliation, every match should come with the rule that made it, the version of that rule, who approved it, and a plain reason your team can check in under a minute.

TL;DR: An explained match shows five things: the rule (fields compared, amount tolerance, date window), the rule version and its approver, a match tier with a written reason instead of a bare score, the match type (1:1, 1:many, many:1, many:many), and an exception trail for what did not match. AI can propose rules. Only approved, deterministic rules should run.

What should an explained match in AI bank reconciliation show?

AI bank reconciliation is the use of software, including machine learning, to pair bank statement lines with ledger entries and route the leftovers to review. The pairing is easy to demo. What survives an audit is the explanation attached to each pair.

Top-ranking pages on this keyword are mostly vendor pages, and they describe matching well. Numeric's guide explains that one-to-one matches rely on amount, posting date and reference, and that complex matches need "adding entries together to match amounts on the other side" (Numeric). HighRadius's product page describes AI that matches "based on historical behavior, amounts, timing, and supporting documentation" (HighRadius). BlackLine's Transaction Matching page describes "applying business rules to automatically match transactions and flag exceptions" (BlackLine). What these public pages rarely show is the record itself: what a reviewer or auditor actually sees next to a matched pair.

Here is the minimum that record should carry:

  • Rule ID, version, approver: which rule fired, which edition, and who signed off on it.
  • Fields compared: for example amount, value date, remittance reference, counterparty.
  • Tolerance and window: the allowed amount difference and date range, as numbers.
  • Match type: 1:1, 1:many, many:1 or many:many, with every line on both sides listed.
  • Tier and reason: which tier matched and one sentence on why, in terms of the fields above.
  • Difference handling: if amounts differ, how much, and which account absorbed it.

A confidence score without a reason is a number you cannot test. A tier plus a reason is something a reviewer can check against the source data.

What does an explained match record look like?

Below is an illustrative record for a customer payment that settles three invoices. The data is invented for this example.

  • Bank line: ACH credit, $18,742.00, value date Oct 6, counterparty "NORTHWIND LOGISTICS LLC", remittance text "INV 3108 3112 3115"
  • Ledger lines: INV-3108 $6,120.00; INV-3112 $9,402.50; INV-3115 $3,219.50 (open AR, total $18,742.00)
  • Match type: 1:many (one bank line to three invoices)
  • Rule: RR-114 v2, "Remittance reference, 1:many", approved by the Assistant Controller on Aug 14
  • Fields compared: invoice numbers parsed from remittance text; sum of open invoice amounts vs bank amount; customer on invoice vs counterparty mapping
  • Tolerance: $0.00 on amount; date window 0 to 10 days after invoice due date
  • Tier: Reference
  • Reason: "All 3 invoice numbers found in remittance text. Invoice total equals bank amount exactly. Counterparty maps to customer C-0412."
  • Difference posted: none

Now the part most tools skip. On the same statement, a $4,980.00 wire arrived with no remittance text. No rule matched it, so it went to exceptions with its own explanation:

  • Exception: EX-2291, wire credit $4,980.00, no reference
  • Closest candidate: INV-3120, $5,000.00, same customer
  • Why it failed: RR-114 requires a parsed reference; RT-031 (tolerance, 1:1) allows $10.00 and the difference is $20.00
  • Status: assigned to AR analyst, open

The exception trail matters as much as the matches. An unmatched item should say which rules it was tested against and which condition failed, so the reviewer starts from the gap instead of from zero.

Which match types should the rules cover?

Every bank reconciliation eventually hits four shapes, and the rule has to say which one it handles:

  • 1:1: one bank line, one ledger entry. A vendor payment clearing one bill.
  • 1:many: one bank line, several ledger entries. One customer payment for three invoices, as above.
  • Many:1: several bank lines, one ledger entry. A card processor settling one day of sales in several deposits.
  • Many:many: several on each side. Batched payroll funding against split journal entries.

The more lines on each side, the more ways a wrong combination can reach the same total. Oracle's Cash Management documentation, for example, states that amount tolerances "can only be used in one to one matching scenarios," and that for one to many, many to one and many to many the amounts must be equal (Oracle). Whether or not your system works that way, the explanation should state the tolerance that applied, so nobody has to guess.

How should match rule tiers work, from exact to model-suggested?

Rules should run in order of strictness, and each tier should carry a different explanation and a different reviewer action.

Table of AI bank reconciliation match rule tiers showing example rule, required explanation and reviewer action for exact, reference, tolerance, model-suggested and exception tiers
  • Tier: Exact; example rule: amount equal, same date, bank reference equals ledger reference; explanation must show: rule ID, version, fields compared; reviewer action: sample-check each period.
  • Tier: Reference; example rule: invoice numbers parsed from remittance, amount equal, 0 to 10 day window; explanation must show: parsed text and which field it came from; reviewer action: review parse failures and partial hits.
  • Tier: Tolerance; example rule: 1:1, amount within $10, date within 3 days; explanation must show: actual difference, limit, account that absorbed it; reviewer action: approve difference postings above a set amount.
  • Tier: Model-suggested; example rule: name similarity plus amount pattern, no rule met; explanation must show: signals used and why no stricter tier matched; reviewer action: accept or reject each one, no auto-post.
  • Tier: Exception; example rule: nothing matched; explanation must show: rules tested and the condition that failed; reviewer action: investigate, resolve, document.

Two design choices keep this honest. A match should come from the strictest tier that fits, and the record should say so. And the model-suggested tier should never post on its own: it is a queue for humans. We made the same argument about where the model should stop in typed decisions vs generated prose.

Yoraito is in early access. Each approved match rule is SQL your team and your auditor can read and rerun. Join the waitlist.

What does a matching rule look like in SQL?

A rule your auditor can read is a rule your auditor can rerun. Here is an illustrative version of the tolerance rule RT-031 from the exception above: 1:1, amount within $10.00, value date within 3 days of the ledger date.

-- RT-031 v3 | Tolerance 1:1 | approved 2026-08-14
SELECT b.bank_line_id, g.gl_entry_id,
b.amount - g.amount AS amount_diff,
b.value_date - g.entry_date AS day_diff
FROM bank_lines b
JOIN gl_cash_entries g
ON g.account_id = b.account_id
AND ABS(b.amount - g.amount) <= 10.00
AND b.value_date BETWEEN g.entry_date - 3 AND g.entry_date + 3
WHERE b.match_status = 'unmatched'
AND g.match_status = 'unmatched';

The limits are literal numbers, not weights inside a model. The output returns the difference, so the explanation can state it. The header names the version and approval date, so an August match traces to the rule text live in August. A production rule also needs a tie-break: if one bank line fits two ledger entries, send both to exceptions with the reason "multiple candidates" rather than picking one silently.

How should an AI-proposed rule become an approved rule?

This is where AI earns its place: spotting patterns across thousands of past matches, like a processor that always settles two days late. Atlar's guide describes a system that "can suggest rules to cover similar scenarios," and users "accept, modify, or dismiss these suggestions" (Atlar). Numeric's guide similarly describes learned matching "within Controller-defined guardrails" (Numeric).

The control question is what happens between suggestion and production. A sound lifecycle looks like this:

Flow showing an AI-proposed bank reconciliation match rule moving through draft, backtest, approval, versioned release and monitoring before it runs
  • Propose: the model drafts a rule in plain terms (fields, tolerance, window, match type) plus the past matches behind it.
  • Translate: the rule becomes deterministic logic, SQL or equivalent, that gives the same result on the same data.
  • Backtest: run it on prior periods and list what it would have matched, including conflicts with approved matches.
  • Approve: a named person with authority signs off, separate from whoever drafted it.
  • Version: the approved text gets an ID and version; old versions stay readable.
  • Monitor: track how often the rule fires and how often reviewers reverse it.

The model can propose a rule; only an approved, versioned, deterministic rule should run. That line is what lets you answer an auditor under AS 1105, which asks auditors to test the accuracy and completeness of information the company produces, or the controls over it. A versioned rule with an approver is a control you can show. A model weight is not. For the wider set of controls finance leaders expect, see why finance teams are still wary of giving AI control, and for the same logic applied to entries, AI journal review needs governance, not just automation.

Frequently asked questions

Can AI fully automate bank reconciliation?

AI can match routine lines and propose new rules, but it should not approve its own rules or post model-suggested matches without review. The dependable setup: deterministic rules for known patterns, human review for suggestions, and an exception queue with reasons for the rest.

What is the difference between a match rule and a confidence score?

A match rule states exact conditions, such as amount equal and date within 3 days, and gives the same answer every time. A confidence score estimates likelihood. Scores are useful for ranking suggestions, but a match should be explained by the rule and reason, not the score alone.

Should amount tolerances apply to one-to-many matches?

Be careful. The more lines involved, the easier it is for a wrong combination to land within tolerance. Some systems, including Oracle Cash Management, allow amount tolerances only for one-to-one matches. If you do allow them, the explanation should state the difference and where it was posted.

How do you keep matching rules audit-ready?

Give each rule an ID, a version, an approver and an approval date. Keep old versions readable, link every match to the version that made it, and log reversals. That lets an auditor rerun a rule on the same data and get the same result.

Where should you start with your own match rules?

Pull last month's reconciliation and pick ten matches at random. For each, can you name the rule, its version, its approver and the reason in one sentence? Then pick five exceptions: can you see which rules they failed and why? Whatever you cannot answer is the gap your auditor will find too. For how AI trust is tracking across finance teams this year, see our 2026 state of trust in finance AI.

Want the rule behind every match? Yoraito shows the rule behind each match and each exception. Join the waitlist for early access.