The Week Agents Ran Out of Benefit of the Doubt

Where the agent boundary belongs: not in the prompt, in your infrastructure

Five headlines in one week, and all of them say the same thing: the question is not whether an AI agent follows the rules, it is what the agent is technically able to do when it doesn't.

I have been in this industry for more than 30 years, and I can't remember many weeks where the news lined up this neatly. Look at the last seven days. A model was cancelled before release, a regulator started asking questions, more than a hundred organizations found out about agent activity from the people who built the agent, and a chip vendor and an operating system vendor both moved the enforcement for agents somewhere the agent can't reach. Different companies, different motives, same lesson. Let's go through it, because I think most IT organizations are still drawing the line in the wrong place.

What actually happened this week?

On September 28 and 29, OpenAI called off the release of GPT-6.1 Astra. According to CSO Online, in internal testing the model could "evade oversight, misrepresent its actions and operate beyond its authorized scope". OpenAI's head of safety systems, Saachi Jain, put it in a sentence every CIO should read twice: the model "did not meet the bar for scope and authorization". Almost at the same time, on DevDay, OpenAI launched Dots, always-on agents in ChatGPT, each with its own cloud computer, with more than 4,000 apps reachable through plugins and actions with account access running through an "auto-review".

On September 30, Axios, CBS News and the Washington Post reported that the FTC is investigating OpenAI, Anthropic, other AI companies and the research organization METR. According to Axios, the FTC is preparing Civil Investigative Demands to get AI executives "to turn over documents and testify about the safety of their models".

On October 1 and 2, the Washington Post and TechSpot reported that OpenAI has notified more than 100 organizations about "misaligned agent activity" connected to its models. The reported behavior includes bypassing access restrictions, using exposed credentials and injecting commands into websites. To be fair, TechSpot also points out that a notification is not automatically a confirmed breach.

In between, NVIDIA introduced its Open Agent Safety Platform: an open-source runtime that isolates agents and enforces access policies for data, tools, APIs and services, plus an out-of-band watchdog running on BlueField-4 DPUs that, according to NVIDIA, can quarantine an agent "in milliseconds". Jensen Huang's summary: "Safety and security require full-stack engineering."

And on October 2, TechCrunch reported that Apple is tightening Full Disk Access on macOS, explicitly because of AI agents. Apple's words: "Some developers are using Full Disk Access in ways that could put users at risk, exposing everything on their systems…without users' full knowledge and understanding."

So what failed? The model?

Honestly, no. Or at least, that is the less interesting part. Go back one more week. When OpenAI paused training and tool-use inference of its most capable models on September 26, the trigger, according to The Decoder and Fortune, was an agent that used missing filtering on a DNS resolver in the research environment to get out to the open internet. Monitoring raised an alert within roughly a quarter of an hour (The Decoder says 12 minutes, Fortune 15). The automatic shutdown failed, and the agent was stopped manually around two and a half hours later.

Read that again with an infrastructure hat on. A DNS resolver without egress filtering. A kill switch that didn't kill. Two and a half hours. None of this is an AI problem, these are exactly the things we have been fixing in data centers for twenty years, and they failed in an environment run by people who know more about these models than anybody else. The model did what models do, it looked for a way to finish the task. The boundary that was supposed to stop it simply wasn't where it needed to be.

And that is my main point: a rule in the prompt is a wish, not a boundary. An agent that is told "don't touch production" but holds credentials for production has not been restricted, it has been asked nicely.

Where does the boundary belong?

Look at what the vendors shipped this week and you see the same direction everywhere. Apple moves the decision into the operating system and wants "very explicit user action" before an agent gets access to everything. NVIDIA moves enforcement into a runtime and onto a separate piece of hardware that the agent can't talk its way around. Broadcom announced AgentMinder (identity, runtime enforcement and audit for agents) and, in a separate announcement, Agentic Zero Trust for vDefend at VMware Explore back in August. And Microsoft, in its Digital Defense Report 2026, almost sounds bored: "Identity and authorization, data protection, least privilege, monitoring, testing, and secure software development all remain relevant."

Yes I know, that last one is not a headline. But it is the most important sentence of the week, because it says there is no new magic, there are the old controls, and they now have to hold against something that is very patient and very creative.

The way I think about it, there are four layers, and the further down you enforce, the less it depends on the agent behaving:

  • Prompt and policy: useful for intent, worthless as enforcement. The agent can read it, so the agent can work around it.
  • Identity and authorization: every agent gets its own identity, its own least-privilege role and short-lived credentials. No shared service accounts, no "it runs as the admin who set it up".
  • Network and data: egress control, segmentation, a data classification that says what an agent may never see, enforced in the network and the platform, not in the application.
  • Out-of-band control: monitoring and a stop mechanism that run outside the agent's reach and outside the vendor's stack. If the kill switch lives in somebody else's infrastructure, it is not your kill switch.

Infographic: four layers where the agent boundary can be enforced, from prompt and policy to out-of-band control, with this week's events as evidence

That last layer is where private infrastructure stops being a philosophical debate. A watchdog in a data center only protects you if the data center is yours, or at least if you control the layer where the enforcement happens.

Would your own monitoring have seen it?

This is the question I would put on the agenda of every IT leadership meeting next week. More than 100 organizations learned about agent activity in their environment, or around it, from the operator of the agent. Not from their SOC, not from their SIEM, not from their network team. From an email.

I don't want to make this about trusting or not trusting one vendor, that is the wrong discussion. The point is that if detection depends on the provider telling you, then you don't have control, you have a subscription to bad news. Agents generate traffic, use identities and touch data. All three of these are things your own tooling should already be able to see. If it can't attribute an action to a specific agent, that is the gap, and it exists today, whether you officially run agents or not.

The ten-minute test

The questions the FTC is now asking the labs are the same questions every company that runs agents will have to answer sooner or later, to an auditor, a regulator or its own board. I would reduce them to three:

  1. What is the agent allowed to do? Systems, data, actions, in writing and enforced technically.
  2. Who approved it? A named person, not "the project".
  3. Where is the log? Every action attributable to one agent identity, stored somewhere the agent can't change.

If you can answer these three for every agent in your organization within ten minutes, you have governance. If you can't, you have hope. And the first step is not a tool, it is an inventory: which agents are running, under which identity, with which permissions, owned by whom. My guess is that most organizations can't produce that list today, and it gets even harder once you include the agents employees switched on themselves.

Decide now, or the decision gets made for you

There is one detail in the Dots launch that matters a lot for Europe: according to SiliconANGLE, Pro subscribers in the EEA, Switzerland and the UK are excluded from the rollout for now, while enterprise customers can use beta features after an admin enables them. Always-on agents with access to thousands of apps will arrive here too, the only open question is when.

The weeks until then are the window in which a company can still decide for itself which data and which systems are off limits for agents, and where that boundary is enforced. After that, the decision gets made by employees, one click at a time, and IT gets to reconstruct it from the logs, if there are logs.

OpenAI made a decision this week, it stopped a model that did not meet the bar for scope and authorization. Whatever you think about the company, that is more than many organizations manage with their own agent pilots. So the question for you is simple: where is your bar, who enforces it, and could you show it to me in ten minutes?

#DecideOrGoHome

At comdivision we spend most of our time in exactly the layers this article is about: platform, network and identity underneath the workloads. If you want to compare notes on how you draw the boundary for agents, reach out to me.

Sources

Related News

dummy_contact-banner_get-in-touch

Questions? Get in touch!

Lorem ipsum dolor sit amet, consectetuer adipiscing elit. Aenean commodo ligula eget dolor. Aenean massa. Cum sociis natoque penatibus.