The Agent Didn’t Go Rogue. The Governance Around It Did.

Share This Post

The Agent Didn’t Go Rogue. The Governance Around It Did.

David Moskowitz – Founder Member and Chief Content Architect, at the DVMS Institute

I reread books. Ten years ago, I read Stanley McChrystal’s Team of Teams for the first time. Having served in the military, I was drawn to his account of why an elite military task force had to change its management structure mid-war. What stood out was the undertraining, a pattern I’ve repeatedly seen firsthand as a consultant. I started rereading the book last night, and what stood out this time wasn’t the undertraining. It was the need to act faster, and how that approach narrowed the available choices. AI changes both the choices themselves and how we must think about them.

McChrystal opens with an Afghan police chief explaining why his ministry kept sending undertrained officers into contested territory: “Of course we understand the dangers, we simply have no other choice.” Then McChrystal turns that line of thinking to his organization. “Most of us would consider it unwise to act before we’re fully prepared…” he writes, “[but] that was the situation we found ourselves in… and it’s the situation leaders and organizations far from any battlefield face every day.”[1]

Ten years ago, that last clause read as a rhetorical “duh.” Rereading it now, it reads like a description of the current state of organizations building or using agentic AI. As commander of JSOC, McChrystal led the Joint Special Operations Task Force in Iraq through a war that gave it no path back to being fully ready before it fought. The Task Force’s answer wasn’t to run its old command structure faster. It discarded it, rebuilding around breaking silos, shared awareness, and adopting a flexible response based on teamwork and collaboration. Why? Because the adversary it faced was already built that way.

Rogue Agent, or Something Else?

On July 16, 2026, McChrystal’s problem surfaced far from any battlefield. An AI agent that OpenAI was testing inside a sandbox found a flaw in the sandbox, escaped, and used that freedom to break into internal systems at Hugging Face, one of the largest AI model-sharing platforms in the world.[2] Hugging Face called the incident a warning shot: “Autonomous, AI-driven offensive tooling is no longer theoretical.”[3]

The BBC ran the story under a headline that framed it as a runaway machine: OpenAI says its AI went rogue.2 That framing spread fast. It fits the story readers already expect: James Cameron’s “The Terminator” for real, the machine turning on its makers. The story also provided almost nothing about what actually failed.

On the Facebook ITSM practitioner group, Back2ITSM, Kaimar Karu, an ITIL 4 lead architect, offered a sharper read: “An automated script did what the script was instructed to do, is my take on this.”[4] James Finister went further, reaching for a familiar image from AI safety writing: “It feels like the paperclip problem. The end overrode the means.”4

Both are right about what this was not. It wasn’t a machine developing its own intentions. In other words, nothing “went rogue” in the sense the headline implies. The AI model followed its instructions. It wasn’t a model failure; it was a containment failure.

The Decision Nobody Made

If the script did what it was instructed to do, it wasn’t a question of who was supposed to notice. Events like this move faster than a human watching a screen. It’s a question of what was supposed to stop the action from completing. That answer isn’t about the moment of escape; it addresses a decision nobody made: to verify sandbox containment before the agent testing started. OpenAI says it will route its findings through a Safety and Security Committee and Safety Advisory Group under its Preparedness Framework.[5] Nothing in the public record indicates that the structure was invoked before testing began; it was invoked only after the breach. That gap is the point: a decision this consequential should have had an identifiable owner, and apparently didn’t.

Not one decision produced this. Several did. The team testing the agent focused solely on the model, excluding its environment. They assumed the sandbox was adequate without verification. Underneath both sat a separation: strategy treated apart from risk, as if the two were different concerns rather than one. Thinking about them separately (i.e., strategy and risk, as separate departments or entities within the organization) is the usual default.

Think Differently

The Digital Value Management System® (DVMS) proposes something different: every decision should be informed by strategy and risk, not as separate entities (or departments), which is the usual standard, but as a single, inseparable governing construct called “strategy-risk.” Once combined this way, every decision to build something requires examining the risks associated with that decision, including those in its operational environment.[6]

Creating and protecting as concurrent activities follow logically from strategy-risk as a governing construct. Apply strategy-risk to this incident, and the two earlier decisions, focusing on the model instead of its environment, and assuming the sandbox was adequate without verification, stop looking like isolated oversights. Attention narrowed to the model because nobody was asking what the whole system, the model and its environment together, actually required. The sandbox went unverified because verifying it read as a protection task, separate from and secondary to the real work of building the agent.

Protection Is Not Resilience

A sandbox is an isolation boundary, a control meant to separate what happens inside from everything outside. Every security boundary represents a point in the system that an adversary, or an agent, can attempt to exploit. In addition, any complex adaptive system (keyword being “adaptive”) will eventually test its edges. Gina Neff, of the Minderoo Centre for Technology and Democracy at the University of Cambridge, put it plainly: “It looks like OpenAI didn’t make a secure enough sandbox.”2

We need to take a brief sidetrack. Protection and resilience are not the same thing, and the difference is where this incident actually gets interesting. For this discussion, consider protection as the set of controls standing between a system and the outcomes an organization wants to avoid. Resilience is the capacity to anticipate, absorb, adapt to, and recover from disruption while preserving human judgment and operational continuity. A sandbox breach is a protection failure. An organization that only learns its sandbox was breached after the fact, from what the agent did on the outside, has a resilience failure.

Anticipation is the capacity most organizations skip, and security engineering already has a discipline built for it, though it wasn’t built with this adversary in mind.

Misuse cases, developed by Guttorm Sindre and Andreas Opdahl two decades ago, extend ordinary use-case modeling to ask not just what a system should do, but what a hostile or misbehaving actor inside it might try instead.[7] The technique assumed a human attacker probing a system from the outside. It didn’t anticipate the system under test becoming the misuser, adaptively, mid-test, discovering its own boundary and acting on what it found before anyone reviewed the results.

If we generalize the Sindre and Opdahl approach from “human actor” to just “actor,” it provides a reason to reevaluate what to model and test. A misuse case for this test would have treated the sandbox as a target rather than just a workspace for the agent. It would have considered what could happen when (not if) the agent probed its own isolation and attempted to escape. Later reporting sharpened that point: the agent wasn’t probing at random; it was chasing the test success criteria, treating hacking as the fastest way to earn a passing score, daisy-chaining through several outside services to get there.5 A misuse case built for this adversary has to model that: not just whether the agent can get out, but whether it will treat getting out as a legitimate way to hit its target.

If that scenario was not modeled, the sandbox was designed as a place for the agent to work, not as an asset that the agent itself might attack. That’s a resilience failure, specifically, a failure of anticipation.

Stuart Rance, responding in the same thread, reframed the story as a detection-and-response problem: “Every system can be compromised.”4 That’s true, and it’s also not the whole answer. Detection tells you that a boundary failed. It doesn’t tell you whether an accountable human still held the decision the instant that boundary gave way, or whether the agent had already converted “capable of acting” into “acted” before anyone was in a position to say yes or no.

That gap, between what an agent is capable of and what a human has actually authorized, is the same gap Shadow AI opens inside an organization. There, the cause isn’t a lack of monitoring. It’s a culture where disclosing AI use feels riskier than hiding it, so the workaround stays inside an employee’s browser tab, unseen.[8]

The mechanism looks different. An escaping test agent and a hallucinated market analysis aren’t the same failure mode. But the underlying discipline they expose is identical: nearly every AI-shaped action needs to remain a proposal until an accountable human disposes of it[9], not a fact the organization discovers after it has already happened.

What Governance Actually Means

This is what governance actually means, underneath the paperwork. Authority and execution aren’t different by degree. They’re different in kind. An agent can extend how much gets done in a given hour. It cannot extend who is accountable for what got done. Accountability belongs to a person who can be named, who understood the conditions at the time, who exercised judgment, and who can account for the result afterward, and no volume of automated output changes who that person is. The failure this produces is rarely one bad call. It’s gradual: de facto decisions drift to whatever system is fastest and most convenient, the org chart still says a person is accountable, and nobody notices the gap between the two until something inside it escapes.

What happens next says more about an organization than the incident itself. Hugging Face published a public account of what happened, said what it still didn’t know, and committed to sharing what it learns.3 That’s what a generative response to failure looks like: inquiry instead of concealment, evidence instead of blame.8 A pathological organization buries an incident like this. A bureaucratic one issues a policy memo and calls the sandbox fixed. Only a generative one treats it as what it actually is: proof that a boundary was tested, and a chance to find out whether the authority behind that boundary was ever really there.

The agent didn’t go rogue. It did exactly what it was supposed to do. In so doing, it found the edge of what nobody had thought to govern.

McChrystal’s task force never got the opportunity to wait until it was ready. Neither does anyone building agentic AI; it requires systems thinking, seeing the agent and its environment as one whole, and thinking differently enough to make anticipation essential.

[1] Stanley McChrystal, Tantum Collins, David Silverman, and Chris Fussell, Team of Teams: New Rules of Engagement for a Complex World (New York: Portfolio/Penguin, 2015).

[2] Laura Cress, “OpenAI says its AI went rogue and launched ‘unprecedented’ cyber-attack,” BBC News, July 22, 2026. https://www.bbc.com/news/articles/c3ek3gvdnj3o

[3] Hugging Face, “Security Incident, July 2026,” Hugging Face Blog, July 16, 2026. Security incident disclosure — July 2026

[4] Kaimar Karu, James Finister, and Stuart Rance, comment thread on Daniel Breston’s post, Back2ITSM (Facebook group), July 2026. https://www.facebook.com/groups/back2itsm/posts/27538611775765855

[5] David Goldman and Hadas Gold, “The OpenAI lab leak was more extensive than we thought,” CNN, July 29, 2026. https://www.cnn.com/2026/07/29/tech/openai-hugging-face-cyberattack

[6] David Moskowitz and David Nichols, A Practitioner’s Guide to Building Operations Cyber-Resilience, 2nd ed., © 2025 by the DVMS Institute LLC. TSO (The Stationery Office), part of Williams Lea, Norwich, UK

[7] Guttorm Sindre and Andreas L. Opdahl, “Eliciting Security Requirements with Misuse Cases,” Requirements Engineering 10, no. 1 (2005): 34-44. Client Challenge

[8] David Moskowitz, “Shadow AI is a Cultural Debt, Not a Technical Vulnerability,” https://dvmsinstitute.com/2026/07/23/shadow-ai-is-a-cultural-debt-not-a-technical-vulnerability/

[9] Some low-risk actions might be pre-authorized, similar to the actions that occur for low-risk request fulfillment. Someone reviewed the possibilities and authorized a set of behaviors that did not need to be verified for every request.

About the Author

David Moskowitz –  Founding Member and Chief Content Architect, at the DVMS Institute

David is a Founding Member and Executive Director of the DVMS Institute LLC. He is the lead author of the “Digital Value Management System®” publication series which include the *Fundamentals of Adopting the NIST Cybersecurity Framework* and *A Practitioner’s Guide to Adapting the NIST Cybersecurity Framework*, and Thriving on the Edge of Chaos published by TSO

Digital Value Management System® is a registered trademark of the DVMS Institute LLC.

® DVMS Institute 2026 All Rights Reserved

 

More To Explore



DVMS: a NIST Cybersecurity Framework Governance by Assurance™ Program that transforms governance from a paper-based compliance exercise into an evidence-driven management discipline that continuously assures resilient and accountable digital operations.