PS HarriJaakkonen:~/Blog/Posts>cat ./microsoft-ai-security-september-2026-part-2.html

When AI Starts Acting: Why Agent Security Is Moving Beyond Prompts

AI agent proposal passing through deterministic policy before execution

For a long time, the worst thing an AI assistant could do was give you a bad answer. Then agents started getting tools. A bad answer could become a bad action.

Imagine the same customer-service team from Part 1, now on a busy afternoon. The refund agent has a browser session, a payment tool and enough authority to finish the job. A document arrives with hidden instructions. The agent reads it, proposes a transfer and reaches the payment screen before anyone sees what happened.

An agent might send email, create an account, modify a database, access a secret, approve a workflow, call an API, execute code or operate a browser. AI security stops being only about securing the model. It becomes about securing authority.

The moment the model gets hands

This is the moment the security story changes. The question is no longer whether the answer is accurate; it is whether the next click, API call or database update is permitted.

Microsoft's 3 September announcement of GPT-6 Astra in Foundry describes a model built to work across applications. The obvious story is the new model. The more interesting security story is what the model can do: reason through tasks, use tools, interact with applications and perform computer-use work.

If a human can click through an internal business system, an agent may increasingly be able to do the same. The absence of an API no longer guarantees the absence of automation. The UI itself can become an execution surface.

An interface becomes an action surfaceAgent: Plans the task; Browser: Interacts with UI; Business app: Exposes operations; Action: Changes real dataCOMPUTER USE / 01An interface becomes an action surface01AgentPlans the task02BrowserInteracts with UI03Business appExposes operations04ActionChanges real dataConceptual flow · Keep execution inside explicit boundaries
Computer use turns familiar user interfaces into an execution surface for an agent.

Prompt injection becomes an authorization problem

Imagine an assistant reading a document containing malicious hidden instructions. If it can only summarize the document, the damage may be limited. Give that assistant tools and the attack path becomes document, prompt injection, agent instruction, tool call and real-world action.

The dangerous question is no longer only whether the model understood malicious content. It is whether the model could turn that content into an authorized action.

I examined the risks around prompts and tool chaining in my earlier CoPhish and Copilot Studio security post. The same design question applies here: which downstream operations remain available when the input cannot be trusted?

Never let the model decide whether an action is authorized

Microsoft Security's 4 September research on securing AI in customer-controlled environments makes an important recommendation: do not let model output directly control privileged operations. Put a deterministic enforcement layer between the model and the action.

The model proposes. Policy authorizes.LLM: Suggests an action; Proposal: Tool + parameters; Policy: Identity / tool / risk; Execute: Only when allowedENFORCEMENT / 02The model proposes. Policy authorizes.01LLMSuggests an action02ProposalTool + parameters03PolicyIdentity / tool / risk04ExecuteOnly when allowedConceptual flow · Keep execution inside explicit boundaries
The model proposes; deterministic policy decides whether execution is allowed.

The model can decide what it wants to do. It should not decide whether it is allowed to do it. A payment, deletion or privileged change needs external rules, approval thresholds and a clear audit trail.

In the refund scenario, the model may propose twenty thousand euros because the customer sounds urgent. A deterministic policy can still reject the amount, require a human checkpoint or limit the agent to creating a review task. The model remains useful, but it is no longer the final authority.

For an example at the tool boundary, my MCP gateway and Zero Trust post describes classifying operations by risk and applying per-tool authorization before execution.

The prompt is not a security boundary

Instructions such as "Never delete production data" and "Only use approved tools" are useful guidance, but they are not security controls. A prompt is advice. An external policy that denies a delete operation against production is a control.

Agent prompt: Never delete production resources.
Policy: DELETE operation against production = DENY

Azure AI Content Safety Prompt Shields can detect adversarial prompts and document content. Use that detection alongside authorization checks on the resulting tool call; the two controls address different stages of the request.

Design as if prompt injection will happen

Microsoft's guidance is healthier when read as a design assumption: prompt injection will happen. Ask what the compromised agent can actually do. Add MFA, Conditional Access, least privilege, segmentation, monitoring, tool restrictions, risk thresholds and human approval.

Limit the impact of prompt injectionInjection: Untrusted input; Identity: Scoped access; Policy: Tool restrictions; Risk: Action thresholds; Human: Approval checkpointDEFENCE IN DEPTH / 03Limit the impact of prompt injection01InjectionUntrusted input02IdentityScoped access03PolicyTool restrictions04RiskAction thresholds05HumanApproval checkpointConceptual controls · apply according to the action and risk
Defense in depth limits what a successful prompt injection can turn into.

The identity controls need to match the caller. My earlier Conditional Access guide for AI agents explains how delegated users, autonomous agent identities and agent user accounts are targeted by different policies.

Computer use makes isolation more important

That is why Microsoft's computer-use guidance belongs in the same story. The browser is not just another interface; it is a room where the agent can see data and press consequential buttons. Give that room scoped credentials, a short lifetime and an activity record.

Computer-use agents may click the wrong button, read unintended information, navigate to an unexpected page or follow malicious content shown inside an interface. Use scoped credentials, approved resources, human checkpoints, activity recording, network restrictions, identity controls and isolated sessions.

A session lasts only as long as the taskTask: Scoped identity; Isolated VM: Restricted session; Logged work: Record activity; Destroy: End the sessionISOLATION / 04A session lasts only as long as the task01TaskScoped identity02Isolated VMRestricted session03Logged workRecord activity04DestroyEnd the sessionConceptual flow · Keep execution inside explicit boundaries
Give computer-use work a short-lived session, scoped identity and recorded activity.

I covered a related execution boundary in Azure Container Apps Sandboxes for agent-generated code. That post looks at sandbox lifecycle, outbound access and the separation between orchestration and untrusted execution.

Runtime trust enters the picture

What if the agent runs on customer hardware, an edge device, remote infrastructure or distributed compute you do not fully control? Before releasing model weights, credentials, customer data, system prompts or retrieval data, you may want proof that the workload is running in an approved environment.

Verify the runtime before releasing secretsRuntime: Requests access; Attestation: Provides evidence; Policy: Verifies conditions; Secrets: Release if approvedRUNTIME TRUST / 05Verify the runtime before releasing secrets01RuntimeRequests access02AttestationProvides evidence03PolicyVerifies conditions04SecretsRelease if approvedConceptual flow · Keep execution inside explicit boundaries
Attestation lets policy decide whether sensitive material can be released to a runtime.

Microsoft's edge AI security research discusses verifying the runtime before releasing sensitive assets. The Azure confidential computing overview provides background on protecting data in use with hardware-based, attested execution environments.

ASCII smuggling shows the crossover

ASCII smuggling uses invisible Unicode characters that may not be obvious to a human reader but still affect how software processes text. An AI system can process hidden instructions even when a person sees only a normal report. The same technique can help traditional phishing evade detections.

What you see is not all the model readsVisible report: Human-readable content; Hidden text: Invisible characters; AI processing: Receives both inputsHIDDEN INPUT / 06What you see is not all the model reads01Visible reportHuman-readable content02Hidden textInvisible characters03AI processingReceives both inputsConceptual flow · Keep execution inside explicit boundaries
Invisible characters can hide instructions from a human while remaining available to software.

AI security techniques and traditional cyber techniques are starting to merge. Attackers will use whichever category works.

I looked at another crossover between conventional application vulnerabilities and AI controls in my analysis of the McKinsey Lilli incident. That case concerned SQL injection and writable system prompts, a different route to influencing a trusted AI application.

Tenant governance becomes AI governance

Microsoft's September Entra update includes Entra Tenant Governance. It helps discover shadow tenants, build governance relationships and detect configuration drift. Shadow tenants can contain copilots, agents, AI applications, MCP connections, enterprise data and external integrations. You cannot secure AI you do not know exists.

For the operational detail, my earlier Tenant Governance GA walkthrough covers related-tenant discovery, governance relationships, configuration baselines and drift monitoring.

The architecture is changing

A traditional AI security model focused on the prompt, a content filter, the model and the response. An agent security model includes agent identity, the LLM, MCP, APIs, other agents, computer use and enterprise data, all behind policy enforcement and approved actions.

Security belongs in the action pathUser: Initiates work; Agent ID: Identifies actor; Agent: LLM / MCP / APIs; Policy: Authorizes action; Action: Approved executionARCHITECTURE / 07Security belongs in the action path01UserInitiates work02Agent IDIdentifies actor03AgentLLM / MCP / APIs04PolicyAuthorizes action05ActionApproved executionConceptual flow · Keep execution inside explicit boundaries
Agent security treats identity and policy as part of the action path, not as decoration around the model.

The new question

AI security used to ask whether the model was safe, whether the prompt was safe and whether the data was safe. Now the most important question may be: what is the agent actually allowed to do?

The strongest enterprise architecture will not assume the model always behaves correctly. It will assume that sometimes it will not, then make sure the system remains secure anyway.

For a broader framework to organize the risks and controls around these systems, see the NIST AI Risk Management Framework.

Sources and further reading