For a long time, the worst thing an AI assistant could do was give you a bad answer. Then agents started getting tools. A bad answer could become a bad action.
Imagine the same customer-service team from Part 1, now on a busy afternoon. The refund agent has a browser session, a payment tool and enough authority to finish the job. A document arrives with hidden instructions. The agent reads it, proposes a transfer and reaches the payment screen before anyone sees what happened.
An agent might send email, create an account, modify a database, access a secret, approve a workflow, call an API, execute code or operate a browser. AI security stops being only about securing the model. It becomes about securing authority.
The moment the model gets hands
This is the moment the security story changes. The question is no longer whether the answer is accurate; it is whether the next click, API call or database update is permitted.
Microsoft's 3 September announcement of GPT-6 Astra in Foundry describes a model built to work across applications. The obvious story is the new model. The more interesting security story is what the model can do: reason through tasks, use tools, interact with applications and perform computer-use work.
If a human can click through an internal business system, an agent may increasingly be able to do the same. The absence of an API no longer guarantees the absence of automation. The UI itself can become an execution surface.
Prompt injection becomes an authorization problem
Imagine an assistant reading a document containing malicious hidden instructions. If it can only summarize the document, the damage may be limited. Give that assistant tools and the attack path becomes document, prompt injection, agent instruction, tool call and real-world action.
The dangerous question is no longer only whether the model understood malicious content. It is whether the model could turn that content into an authorized action.
I examined the risks around prompts and tool chaining in my earlier CoPhish and Copilot Studio security post. The same design question applies here: which downstream operations remain available when the input cannot be trusted?
Never let the model decide whether an action is authorized
Microsoft Security's 4 September research on securing AI in customer-controlled environments makes an important recommendation: do not let model output directly control privileged operations. Put a deterministic enforcement layer between the model and the action.
The model can decide what it wants to do. It should not decide whether it is allowed to do it. A payment, deletion or privileged change needs external rules, approval thresholds and a clear audit trail.
In the refund scenario, the model may propose twenty thousand euros because the customer sounds urgent. A deterministic policy can still reject the amount, require a human checkpoint or limit the agent to creating a review task. The model remains useful, but it is no longer the final authority.
For an example at the tool boundary, my MCP gateway and Zero Trust post describes classifying operations by risk and applying per-tool authorization before execution.
The prompt is not a security boundary
Instructions such as "Never delete production data" and "Only use approved tools" are useful guidance, but they are not security controls. A prompt is advice. An external policy that denies a delete operation against production is a control.
Agent prompt: Never delete production resources.
Policy: DELETE operation against production = DENY
Azure AI Content Safety Prompt Shields can detect adversarial prompts and document content. Use that detection alongside authorization checks on the resulting tool call; the two controls address different stages of the request.
Design as if prompt injection will happen
Microsoft's guidance is healthier when read as a design assumption: prompt injection will happen. Ask what the compromised agent can actually do. Add MFA, Conditional Access, least privilege, segmentation, monitoring, tool restrictions, risk thresholds and human approval.
The identity controls need to match the caller. My earlier Conditional Access guide for AI agents explains how delegated users, autonomous agent identities and agent user accounts are targeted by different policies.
Computer use makes isolation more important
That is why Microsoft's computer-use guidance belongs in the same story. The browser is not just another interface; it is a room where the agent can see data and press consequential buttons. Give that room scoped credentials, a short lifetime and an activity record.
Computer-use agents may click the wrong button, read unintended information, navigate to an unexpected page or follow malicious content shown inside an interface. Use scoped credentials, approved resources, human checkpoints, activity recording, network restrictions, identity controls and isolated sessions.
I covered a related execution boundary in Azure Container Apps Sandboxes for agent-generated code. That post looks at sandbox lifecycle, outbound access and the separation between orchestration and untrusted execution.
Runtime trust enters the picture
What if the agent runs on customer hardware, an edge device, remote infrastructure or distributed compute you do not fully control? Before releasing model weights, credentials, customer data, system prompts or retrieval data, you may want proof that the workload is running in an approved environment.
Microsoft's edge AI security research discusses verifying the runtime before releasing sensitive assets. The Azure confidential computing overview provides background on protecting data in use with hardware-based, attested execution environments.
ASCII smuggling shows the crossover
ASCII smuggling uses invisible Unicode characters that may not be obvious to a human reader but still affect how software processes text. An AI system can process hidden instructions even when a person sees only a normal report. The same technique can help traditional phishing evade detections.
AI security techniques and traditional cyber techniques are starting to merge. Attackers will use whichever category works.
I looked at another crossover between conventional application vulnerabilities and AI controls in my analysis of the McKinsey Lilli incident. That case concerned SQL injection and writable system prompts, a different route to influencing a trusted AI application.
Tenant governance becomes AI governance
Microsoft's September Entra update includes Entra Tenant Governance. It helps discover shadow tenants, build governance relationships and detect configuration drift. Shadow tenants can contain copilots, agents, AI applications, MCP connections, enterprise data and external integrations. You cannot secure AI you do not know exists.
For the operational detail, my earlier Tenant Governance GA walkthrough covers related-tenant discovery, governance relationships, configuration baselines and drift monitoring.
The architecture is changing
A traditional AI security model focused on the prompt, a content filter, the model and the response. An agent security model includes agent identity, the LLM, MCP, APIs, other agents, computer use and enterprise data, all behind policy enforcement and approved actions.
The new question
AI security used to ask whether the model was safe, whether the prompt was safe and whether the data was safe. Now the most important question may be: what is the agent actually allowed to do?
The strongest enterprise architecture will not assume the model always behaves correctly. It will assume that sometimes it will not, then make sure the system remains secure anyway.
For a broader framework to organize the risks and controls around these systems, see the NIST AI Risk Management Framework.
Sources and further reading
- GPT-6 Astra: frontier intelligence for work in Microsoft Foundry
- How to secure edge AI in customer-owned environments
- ASCII smuggling crosses over from AI prompt injection to phishing evasion
- Agent identity, MCP and A2A authentication
- Hosted agents and VM-isolated sandboxes
- What's new in Microsoft Entra, September 2026
- Azure AI Content Safety prompt shield concepts
- Azure confidential computing overview
- NIST AI Risk Management Framework


