Insights
← All Insights
AI GOVERNANCE

Governance that speeds AI adoption

Article 6 of 6
EXECUTIVE SUMMARY
  • Governance done well speeds AI adoption. Ambiguity, not control, is what stalls finance teams.
  • The governance floor starts with finding what is already in use, approving the tools worth keeping and recording what each one may do.
  • At the general ledger level, AI has read only access by default and every posting needs human approval, with a narrow tested exception for proven batches of identical matches.
  • Controls tighten with sensitivity and autonomy. Decide where sensitive data and models may sit, and give standing agents their own identity, control record and harness.
  • A named group keeps the governance framework current through a light monthly review: the tool log, decision rights, human review, forecasts and the evidence trail.
  • The rules already apply, from UK GDPR to the EU AI Act's near term duties, and accountability stays with a named individual, never the tool.
Read the full article ↓

The minimum governance framework to put in place now, the full framework and its ownership, and the routine that keeps it current, from model updates to the new rulebooks.

Governance has a reputation as a brake. For finance teams adopting AI, that reputation is rarely deserved. The teams that generally progress fastest put a few clear rules in early. Ambiguity is usually the greater obstacle. When nobody knows what a tool may touch, who owns its output or where a person must check the work, progress stalls.

The starting point is a short initial governance policy, as outlined in the strategy article, which an LLM can help draft from agreed principles. That policy provides the minimum governance needed while AI use is limited to a few assistants handling routine tasks under close review. As assistants give way to more complex automated and agentic workflows, the policy has to develop into a full governance framework.

Both frameworks depend on clear responsibility. IT owns security. The privacy officer oversees personal data. Finance sets the rules for anything that touches the numbers and remains accountable for those numbers whatever the tool does. This article covers the key steps needed to set the governance floor and build the full framework. It also explains the monthly review that keeps the framework current and the new rulebooks finance needs to understand.

The minimum governance framework

Setting the governance floor starts with finding what is already in use. Ask the team directly and work with IT to identify AI applications installed on company devices or used through a browser. Expense claims are also worth checking for AI licences, a task AI itself can help with.

The search will usually find something. A Smartsheet vendor survey published in January 2026 found that 70 percent of 1,550 operations professionals reported AI use outside company policy. Finance is unlikely to be different. The review should include anything the team has built with AI as well as the tools it has bought.

Tools worth keeping need approval before further use. The tool log records what each one does, the systems and data it reaches, whether it reads or writes, who has access and the vendor's security certifications. Approval also looks ahead. Before the function comes to depend on a tool, confirm that its data and activity logs can be exported and that the contract allows a clean exit. The same approval applies to anything the team builds. A script or small application created with an LLM needs an owner, a record of the data it touches and a way to switch it off. Once its output feeds the ledger or a reported figure, it is held to the same controls as a bought tool.

That approval also settles three operating questions for every tool. It records what may run unsupervised, where a person must review the work and how the output is tested before anyone relies on it. At the general ledger level, one rule is fixed from day one. Read access is the default and a person approves every posting, even one that appears to be a perfect match. Automatic posting can follow only for proven batches of identical matches, where the judgement sits in the tested matching rule rather than each posting. A typical example is a batch of cash receipts matched to invoices on the same reference and amount.

Access to systems otherwise follows the person. An AI assistant used by an analyst has the analyst's permissions and nothing more. A standing agent is different. It starts work on a schedule or trigger, such as overnight invoice chasing, rather than waiting for a person to ask. It therefore needs its own identity, permissions and limits, which are covered in the full governance framework.

The data each tool may receive is also decided individually. While the team is learning, the simplest protection is to exclude confidential data and use material that would cause no harm if it leaked.

Where confidential data is needed, the minimum protection is a business or enterprise tier licence, whose terms exclude customer content from model training by default. Work stays on company managed devices, keeping personal accounts and unapproved applications away from company data. Written retention rules state how long prompts, outputs and logs are kept and when they are deleted.

Some businesses go further and keep both data and models entirely within their own environment. Whether that is necessary depends on the sensitivity of the data, the organisation's obligations and its tolerance for vendor access. The available options in this respect are considered next.

The governance floor applies to every use case. It remains sufficient while a person directs the work, the tool stays within that person's existing permissions and the output is closely reviewed. The full framework adds process controls when the use case involves sensitive data, access to a live system or material financial consequences. It adds standing agent controls when the tool starts work independently.

The full governance framework

One principle governs the full framework. Controls tighten as data becomes more sensitive, tools act more autonomously and actions become harder to reverse. Data is where that principle usually bites first. Keeping confidential material out protects the governance floor, but it also restricts the work AI can do. Most functions therefore allow approved tools to use confidential data under the Anthropic and OpenAI terms that exclude customer content from training.

However, model training is not the only exposure. The Sunday Times reported in July 2026 that Anthropic and OpenAI both monitor how their tools are used. Anthropic also acknowledged that it studies aggregated, de identified usage to improve its products.

For many companies, that residual exposure will exceed what they are prepared to accept. Over time, patterns in prompts, tool use and corrections can also allow institutional knowledge to seep gradually into the AI provider's models and products, even where no single item identifies its source. Competitors may then benefit from capabilities shaped by knowledge they could not otherwise access. The organisation therefore needs one deliberate decision about where its most sensitive data and models may sit, rather than leaving that judgement to each employee and each document.

There are two main ways to keep access under the company's control. The first is a ring fenced cloud tenancy in which the company controls access, retention and encryption. Prompts, outputs and any model tuned on company data remain inside the approved environment under the agreed terms. Customer managed keys strengthen control over encryption, but do not by themselves make the content unreadable to the cloud provider. Where the cloud provider itself must be unable to see the content, the design needs client side encryption or an equivalent confidential computing arrangement.

The second option is an open source model downloaded and run on the company's own infrastructure, so no data needs to go to an outside model provider. Model charges can be far lower than for leading commercial models, and the capability gap has also narrowed. However, lower model charges come with greater operating responsibility. The company supplies the computing capacity and security, maintains the model, tests its quality and supports the service. The strongest commercial models may still perform the hardest judgement work better.

For many small to mid market finance teams, the sensitive core is customer records, deal terms and board papers rather than proprietary product design. A well configured ring fenced tenancy will usually provide enough protection. Running an open model in house becomes more compelling where sensitivity, data residency or cost at scale demands it.

After deciding where the data and model sit, the next question is what AI may change in live finance systems. Write rules are agreed in advance, and changes to the chart of accounts or master data always require sign off. Automatic posting is permitted only for narrow, repeatable cases where the matching rule has been tested and proved reliable. Qualifying batches can then post automatically, every failed match goes to a person, and the exceptions are cleared daily. Every other posting remains subject to human approval. The same standard applies to payments, supplier bank details and payroll, where an error is costly and difficult to reverse.

Those controls must leave a complete evidence trail. For any AI output or action affecting a reconciliation, posting, payment or reported figure, finance must be able to reconstruct what happened from source data to approval. The record identifies the approved use case, source data, instructions, model version, action taken, any exception or override, and the person who reviewed or approved it. This evidence supports management review, internal control testing, internal audit and external audit, whether a person or a standing agent performs the work.

The control model changes when a standing agent starts the work. When a person starts the work, the tool uses that person's permissions and remains under close review. A standing agent starts independently, so it needs its own identity and one agent control record before it can act.

Each standing agent has one control record. It lists the systems the agent may reach, whether it reads or writes in each, and the actions it may take alone. It also records what an AR agent may send to a customer, what a close agent may submit as a draft journal and what a forecast model may publish. Numerical thresholds sit beside named escalation roles. A journal above a set value or a payment above a set limit stops for approval.

The same record assigns accountability. If a controller approves an agent drafted journal, the controller owns that decision. Automatic posting of a proven batch does not break separation of duties. A person has already approved the matching rule, and every failed match still goes to a person.

The harness is the technical layer coded round the LLM that enforces the agent control record. It contains the instructions that define the agent's job and boundaries, together with its identity, permissions and approved connectors. It sets limits on values, volumes and spend, places approval gates and provides the stop switch. The record states what the agent is allowed to do. The harness makes those rules operate and writes the activity log showing what the agent actually did.

The agent control record also keeps the approved configuration, version history, test evidence and any compensating controls. Any change to the instructions, connectors, permissions, limits or underlying model receives normal system change discipline, whoever makes it. Where the technology permits, a new version runs alongside the current version on the same work until its results are proven. The current version remains available if the replacement misbehaves. If fixed versions or parallel testing are unavailable, the record identifies the gap, the compensating controls and the person accepting the risk.

For most small and mid market teams, the simplest route is to buy an agent platform with the main harness controls built in and configure only the finance rules. Other teams assemble those controls on a workflow platform such as n8n or write them in code. The governance standard does not change. Anything assembled or coded internally becomes another system owned by the team. It needs a named owner, its own testing and maintenance when a connected application changes.

The framework needs clear ownership as well as clear controls. A small named group runs it, with finance and IT sharing responsibility for agents and access while finance owns the ledger rules. The group approves use cases and exceptions, reviews evidence and incidents, and escalates material issues through the existing risk and control matrix. It also runs the monthly review and reports to the audit committee or its equivalent. Where the organisation has an internal audit function, that function should periodically test the design and operation of material AI controls independently of the team that runs them. Citi's then CFO Mark Mason described the right posture in Fortune in November 2025: enthusiasm tempered by caution, robust controls, subject matter expertise and dual processes for financial reporting.

Keeping the framework current

The named group keeps the framework current through a light monthly review built around five core checks:

  • The tool log: confirm that it remains current, covers every approved tool, internally built application and standing agent, and names the owner of each. This is the same inventory discipline used in the US Treasury framework.
  • Decision rights and access: confirm that permissions, approval points, numerical limits and escalation roles remain current and are being followed. Check that access remains appropriate and every recorded exception is still justified.
  • Human review and validation: confirm that approval remains in place wherever the decision rights require it. This includes AI output used to explain a reconciliation difference, create a posting or support a reported figure. The output must be explainable, tested and traceable before it reaches the numbers.
  • Forecasts: treat an AI forecast like any other model and have it reviewed by someone other than its builder.
  • Evidence trail: sample material outputs and transactions. Confirm that the approved use case, source data, instructions, model version, action, exception or override, and approval can all be reconstructed from the records. For a transaction, the trail must connect the source record to the posting or payment and the person who approved it. The evidence must support management review, internal control testing, internal audit and external audit.

These five checks reflect a basic difference between AI and conventional system controls. A deterministic system gives the same answer to the same inputs, although bad data or configuration can still make that answer wrong. A probabilistic tool can return a different answer on another run. Validation, traceability and monitoring therefore have to be designed into the first use case rather than added after the fifth.

The monthly review does not replace oversight of a write capable agent while it runs. The harness enforces limits and approval gates continuously. The owner also takes a weekly sample from the activity log and tests both behaviour and transaction traceability. The sample covers failed actions, near misses, unusual volumes, repeated retries, overrides and movements outside approved thresholds. For material transactions, the reviewer also reconstructs the source data, model version, action and approval.

The review also looks for drift, meaning a gradual change in behaviour as the model, instructions or underlying data changes. The stop switch is tested by using it, confirming that the agent halts and a person picks up the work in flight.

Before a write capable agent goes live, the use case owner agrees the failure response and leads it when required. The incident record names who is informed, what is disabled, how affected outputs are quarantined, how the work continues and who authorises restart. It also identifies the affected transactions, decisions, remediation and evidence needed before the agent returns to service. Failures and near misses are tracked until closure, and any use case that does not earn its place is stopped.

The same operating review also covers cost. Agent work is metered, so spending rises with activity rather than headcount. Each agent or tool has a named owner, a spending cap and a place in the benefits ledger, where actual cost is compared with the savings it funds. CFO.com reported in May 2026 that Microsoft wound down internal Claude Code licences after token costs climbed sharply. A bill that even a division of Microsoft reportedly stopped paying is a warning to every smaller finance function.

Once a quarter, the audit committee or its equivalent receives a concise summary of the approved AI portfolio. Drawn from the tool log, agent control records, incident records and benefits ledger, it covers the tools and agents in use, material changes, control or evidence gaps, incidents and near misses, spend against value and any decisions required. The deployment choice is revisited when costs rise at scale, residency requirements change or the function moves from buying to building. The diagram below brings the framework together by showing the governance floor, the additional process and agent controls, and where the common evidence and reporting checks sit.

Layer
What it holds
Agent controls, for standing agents
Its own identity, an agent control record, a harness coded round the LLM, and hard limits on values and volumes.
Process controls, as autonomy and consequence grow
Write rules agreed in advance, tested matching rules before any automatic posting, every failed match to a person, and sign off on master data changes.
The governance floor, from day one
An approved tool log, one rule for confidential information, read only access as the default, a person in the loop on anything that posts, and a named owner for every automated task.
Evidence and reporting, across every layer
A complete evidence trail from source data to approval, the monthly review, the quarterly summary to the audit committee, and the benefits ledger with spending caps.

The rules that already apply

Any governance framework naturally sits inside the law. UK GDPR, confidentiality obligations and contract terms apply to an AI tool as they do to any other system. A new use of personal data may still require its own data protection assessment.

Some financial firms face stricter duties for material third party arrangements. DORA applies to defined categories of financial entity in the EU and governs their ICT third party risk. FCA and PRA rules cover operational resilience and outsourcing for UK firms within scope. Outsourcing a system never outsources accountability.

Three newer rulebooks also matter, although they serve different purposes. The FRC guidance published on 30 March 2026 is written for audit firms rather than finance teams, but it shows the questions an auditor is likely to ask. COSO's February 2026 guidance on internal control over generative AI points the same way: treat GenAI outputs as claims requiring validation, and first ask whether deterministic automation or traditional machine learning would do the job more reliably. The FRC's central principle extends beyond audit: accountability stays with a named individual, never with the tool.

The US Treasury framework published in February 2026 is non binding. Even so, its 230 control objectives, developed with more than 100 financial institutions, provide a practical checklist for building the full framework.

The EU AI Act is law, and June 2026 changed its timetable. It can apply to a UK company where a system is placed on the EU market or its output is used in the EU. The formally adopted delay moves obligations for stand alone high risk systems to 2 December 2027. Transparency duties on chatbots and synthetic content take effect on 2 August 2026 for new systems. Certain existing generative systems have until 2 December 2026 for the content marking and detection obligation.

The Act's AI literacy duty already applies, and role specific training is the obvious way to meet it. The dates do not all move together, and the near term duties are the ones to plan around.

Light while learning, tight where it writes

The evidence kept from the outset is what allows the function to scale later. The tool log, approval and agent control records show what exists, who owns it and what each tool or agent may do. Version and change records show what changed. Activity logs, transaction trails, exceptions and approvals reconstruct each material output or action from the approved use case and source data to the final human decision.

KPMG's May 2026 Global AI in Finance report surveyed 1,013 finance leaders. Organisations able to produce AI audit evidence efficiently reported significant improvement at three to six times the rate of the rest. On error reduction, the figures were 33 percent against 6 percent.

The brake, once again, is not control but ambiguity and absent accountability. Where a second pair of hands would help, aicfopartner.ai does this work, so book a call.

© 2026 aicfopartner.ai
tim@aicfopartner.ai