Before an Agent Writes to Your Systems

How to Govern AI That Acts

For the past two years, most enterprise AI has been conversational. A copilot drafts an email. A search tool finds the policy document. A summary turns a long thread into five lines. If the answer is wrong, a person reads it, frowns and moves on. The damage stays on the screen.

That is changing. The questions operations teams now ask of AI sound different: can the agent raise the purchase order, reschedule the maintenance job, update the supplier record in SAP, close the work order in Dataverse. These are agents that act. When one of them is wrong, the error lands in a system of record, flows into the next process, and may reach a supplier, a customer or an auditor before anyone notices.

The technology to do this exists today on Microsoft. The harder question is what has to be true before you let it. What follows are the six controls I would want in place before any agent writes to a system that matters, and the questions I would ask of anyone, including us, who proposes to build one.


Define the actions, and give the agent nothing else

The first control is the simplest to state and the one most often skipped. An agent should only be able to call specific, named actions that someone has designed and approved. “Update the delivery date on a purchase order line”, with typed inputs: an order number, a line number, a date within a permitted range. “Reassign a work order”, with the technician drawn from a list the system already holds.

What the agent should never have is open access: a database connection it can send raw SQL through, or a general REST client pointed at your ERP with a powerful service account behind it. Open access turns every clever prompt into a potential change nobody designed. Typed actions turn the agent’s freedom into a menu you wrote.

OWASP names this risk Excessive Agency in its Top 10 for LLM Applications, and its first mitigation is to limit what an agent can call “to only the minimum necessary”, avoiding open-ended tools such as a shell command. The UK NCSC’s guidelines for secure AI system development, written with the US CISA and international partners, likewise ask for restricted actions and least privilege wherever a system can act on other systems.

This has a useful side effect. A list of approved actions is something your architects, your security team and your process owners can read, challenge and sign off. It becomes a governance document in its own right, and it makes the scope of the agent visible to people who will never read a prompt.


Set thresholds, and a named person signs off above them

Actions carry different risks. Moving a delivery date by a day is a small thing. Cancelling an order is a bigger one. Reordering stock worth a few hundred pounds sits in a different category from committing a six-figure spend.

The business should set the limits, in its own terms: value, quantity, which suppliers, which sites, which record types. Within those limits the agent can proceed. Anything above them stops and waits for a person, which is the human approval OWASP recommends for high-impact actions. The thresholds belong to the process owner, sit in configuration where they can be reviewed, and change through the same approval route as any other control.

Start the limits low. Raising a threshold after a quarter of clean records is an easy conversation. Explaining why it was set high on day one is a hard one.

When an action waits, it should wait for someone in particular. The approval request goes to a named person, or a named role with a named deputy, in the place they already work. On Microsoft that usually means Teams: a card that shows what the agent proposes, why, and the data it relied on, with approve and decline. Microsoft documents this pattern: a Power Automate approval can be answered in Teams, Outlook or the action centre while the flow waits.

The person approving must be verified. Their identity comes from Entra ID, under the same conditional access and multi-factor rules as the rest of your estate, so every approval traces to an individual and cannot be clicked through by whoever happens to see the message.

This is a gate, and we work to the same idea in our own delivery. Talastron Kinetic AI, our delivery pipeline, runs eight specialist agents with five human gates between them: points where the work stops for review before it moves on. An agent that acts in your operations deserves at least the same discipline.

A gate is only useful if it really stops the work, so the design has to say what happens when nobody answers. The request expires, escalates to the deputy, or falls back to the manual process. It should never proceed by default.

For AI systems classed as high-risk, the EU AI Act makes this law: Article 14 requires that people can oversee such a system while it is in use, override or reverse its output, and halt it “in a safe state”. Whether an agent is high-risk depends on what it does and where it is used, but the article is a useful benchmark for any agent that acts.


Record everything, in your own tenant

Every action should leave a record a stranger could follow: the proposal the agent made, the data it used to reach it, the thresholds it was checked against, who approved or declined it and when, and the change that was finally written, with the values before and after.

Those records belong in your own Microsoft tenant, under your retention policies and your access controls, where your auditors and your security team can reach them without asking a supplier. They should be tamper-evident: written once, protected from quiet editing, and checked so that any alteration shows. When internal audit asks why a record changed on a particular Tuesday, the answer should take minutes to find. The NCSC guidelines ask for the same: log what goes into an AI system, within data protection rules, so that audit and investigation are possible.

A good record also teaches. Reviewing the declined proposals each month shows where the agent’s judgement and your people’s judgement differ, and that is exactly where the actions and the thresholds need work.


Know how every change is undone

Before an action goes live, someone should be able to say how it is reversed. Some changes are easy to undo: a date can be set back, an assignment moved. Others cannot be recalled once they leave the building: a purchase order sent to a supplier, a payment released, a message to a customer. For those, the control is compensation: a defined follow-up action, such as a cancellation or a correcting entry, with its own owner and its own approval.

Ask for this per action, written into the action’s design. If the answer to “how do we undo this?” is “we would have to work it out”, the action is not ready to be given to an agent.


Keep it in your tenant, in the region you choose

An agent that writes to your systems of record should run inside your own Microsoft tenant, deployed to the Azure region you choose, such as UK South, under identities you manage. Your data stays where your policies say it lives, and you can switch the whole thing off without anyone else’s permission. The NIST AI Risk Management Framework asks for the same: clear responsibility and a means to disengage or deactivate an AI system that behaves outside its intended use.

Be precise about isolation, because the words get blurred. Some parts of a solution can be network-isolated: the model endpoint, the storage and the functions that call your ERP can sit behind private endpoints on your own virtual network, with no public route in. Other parts are tenant-isolated: Teams, Entra ID and many Microsoft 365 and Power Platform services run as shared Microsoft services, separated by tenant and reached over Microsoft’s network. Both can be the right choice. Your architecture diagram should say which is which, and your security team should have agreed it.


Questions to ask before you let an agent act

Before any agent is allowed to write to a system of record, I would want a clear answer to each of these.

  1. What exact actions can the agent call, and who approved that list?
  2. Can the agent reach any system by any route other than those actions?
  3. What are the thresholds, who owns them, and how are they changed?
  4. Who approves anything above a threshold, how is their identity verified, and what happens if they do not respond?
  5. What is recorded for each action, where is it kept, and how would we know if a record had been altered?
  6. How is each action reversed, or compensated if it cannot be reversed?
  7. Which tenant and region does it run in, and which parts are network-isolated and which are tenant-isolated?
  8. How do we pause the agent, and who is allowed to?
  9. Who reviews declined and reversed actions, and how often?

If any of these has no answer yet, the agent is not ready to write. It can still propose, and a person can still act on the proposal. That is often the right place to start.


These are the design principles we hold to for every agent we build that can change a system in a customer’s business. Our Trust Centre sets out the certifications we hold today and the work under way, at talastron.com/trust. If you are weighing up where an agent could safely act in your own operations, we would be glad to talk it through: talastron.com/contact.

Sources

  1. OWASP Gen AI Security Project. “LLM06:2025 Excessive Agency.” OWASP Top 10 for LLM Applications, 2025. genai.owasp.org/llmrisk/llm062025-excessive-agency
  2. UK National Cyber Security Centre, US CISA and international partners. “Guidelines for secure AI system development.” November 2023. ncsc.gov.uk/collection/guidelines-secure-ai-system-development
  3. European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 14 “Human oversight.” June 2024. eur-lex.europa.eu/eli/reg/2024/1689/oj
  4. NIST. “Artificial Intelligence Risk Management Framework (AI RMF 1.0).” NIST AI 100-1, January 2023. doi.org/10.6028/NIST.AI.100-1
  5. Microsoft Learn. “Get started with Power Automate approvals.” 2026. learn.microsoft.com/power-automate/get-started-approvals

Not sure whereto start?

Tell us where you are today, and we’ll recommend the right starting point.