Shining a Light on Shadow AI
David Nichols – Co-Founder and Executive Director of the DVMS Institute
Say the words Shadow AI in front of a leadership team and watch what happens to the room.
Someone reaches for a word like exposure. Someone else says containment. The tone shifts, as it does during a data breach or a fraud investigation, when something happens in the dark that the organization is only now, belatedly, discovering. Shadow AI is discussed as a hidden population of rule-breakers, quietly undermining governance from inside the walls, waiting to be found and shut down.
That reaction is not irrational. It is also almost entirely wrong, and the word doing the damage is shadow itself. Shadow implies concealment, and concealment implies intent. It casts ordinary, explainable behavior in the language of a threat, and once behavior is cast as a threat, the only response that feels adequate is a harder one. Detection. Blocking. Consequences. Treat it the way security treats any unauthorized access, because that is what the word shadow tells a leader to feel.
Turn on the actual light, and the monster is mostly not there. What is there, in its place, is something far less dramatic and far more solvable. Most of what gets called Shadow AI is a capability gap wearing a scarier name, and understanding that difference changes what a leader should do about it.
The Belief Does Not Survive Contact With the Reason People Do This
Ask why someone opened a personal AI account instead of using the approved one, and the answer is rarely defiance. It is usually a two-hour task that became a ten-minute task once they found a tool that worked. The employee was not trying to evade the organization. The employee was trying to finish an assignment, meet a deadline, or improve the quality of something under time pressure, which made the sanctioned path feel too slow, too limited, or simply absent for the task at hand.
That distinction matters because it points to a different diagnosis. If the behavior were driven by a desire to break rules, tighter enforcement would work. If an unmet operational need drives the behavior, tighter enforcement does not remove the need. It only removes visibility into how people are meeting it.
This is the same dynamic that produced Shadow IT a decade earlier. Employees adopted unauthorized file-sharing tools, messaging apps, and cloud services not because they preferred to operate outside the rules, but because the approved alternatives were slower, harder to access, or ill-suited to the work. The shadow was never created by employee preference alone. It was created by the gap between what people needed to do their jobs and what the organization provided.
AI did not invent this pattern. It intensified it. A generative AI tool requires no installation, often costs nothing, and is just a browser tab away. It provides immediate leverage for drafting, summarizing, analyzing, coding, and comparing.
An employee who has already experienced that leverage will not view a blanket prohibition as neutral governance. They will see it as an obstacle to the outcome for which they are accountable.
Two Systems Are Running at Once, and Only One of Them Gets Reported Upward
Every organization operates two systems simultaneously.
The first is the paper system. It includes policies, approved tool lists, training modules, and monitoring dashboards. It is carefully built and genuinely useful because it encodes real judgment about what should happen. It is also structurally disconnected from the daily friction that determines what actually happens.
The second is the lived system. The real workarounds, the real time pressure, the real conversations in which someone says, “just use the other one, it is faster,” because it is true. This is where organizational behavior actually lives, in trade-offs that the paper system was never designed to capture.
Senior leaders see the paper system because that is what gets summarized and sent upward through a dashboard or a compliance report. The organization itself operates in the lived system, at the level where the actual work gets done. Those are not the same view, and the distance between them is exactly where risk accumulates unnoticed.
The divide between the paper system and the lived system is nothing new, and it was not invented for AI governance. It appears that an institution relies on documentation to represent reality rather than continuously testing whether the documentation still matches reality. Firms routinely appear fully compliant right up until the week something fails, because every box was checked and nobody was checking what actually mattered. The record and the underlying reality are two different objects, built by different processes, and an organization that only inspects the record has no reliable way of knowing how far reality has drifted from it.
Shadow AI is what that drift looks like when the tool in question is generative AI. The paper system prohibits or tightly controls usage. The lived system shows teams relying on it daily to meet deadlines that the paper system’s approved tools cannot help them meet. Governance that reads only the paper system will never see the gap it is supposed to govern.
The Real Question Is Not How Much AI Is Used
Once you accept that the gap is real, the next mistake is measuring the wrong thing to close it.
A usage dashboard can show logins, prompts, and adoption rates for approved tools. That data helps with cost and capacity planning. It does not explain why unauthorized use exists because it only measures what happened inside the boundary you already control. It cannot see what happened outside that boundary, and what happened outside it is the entire signal you need.
The more useful measure is the gap between what someone needed to accomplish and what the organization actually enabled them to accomplish safely. That gap can be observed directly through a small set of concrete questions.
Was an approved capability available when the person needed it? Could they complete the task without leaving the sanctioned environment? Did the result meet the quality and safety bar that the organization actually cares about? How much friction stood between the need and access? Did the review requirements match the actual consequence of the task, or were they the same regardless of the stakes?
Answer those honestly for a specific team and a specific recurring task, and you will usually find the same pattern. The organization is not failing to control AI use so much as it is taxing legitimate use, which leads people to route around the tax.
That tax framing explains something counterintuitive. When an organization imposes excessive delay and administrative friction on a legitimate need, it does not reduce that activity. It gets the same activity, now harder to see, harder to assure, and far less likely to yield anything the organization can learn from. Suppressing visibility is not the same as control. It is usually the opposite of control, dressed up as discipline.
Govern the Outcome, Not the Tool
If the tool is not the right unit of governance, something else must be, and that is the outcome the work was supposed to produce.
The question that matters is not how much AI was involved in the abstract. It is whether the resulting work advanced a declared outcome within acceptable constraints and with sufficient evidence to support real confidence in the result. That reframing does real work by letting the organization calibrate proportionally rather than uniformly.
Drafting an internal meeting summary and preparing a regulatory filing do not warrant the same level of control. Brainstorming a list of product names and recommending a clinical intervention do not carry anything like the same risk. A single blanket policy, whether it prohibits everything or permits everything, ignores differences that clearly matter to anyone doing the actual work.
Outcome-centered governance instead asks what must remain true for a given outcome class, what constraints apply, what evidence is required, and who has the authority to accept or reject the result. Within that structure, teams get broad access where the outcome and its constraints allow, and tighter review where the consequences of being wrong are high. The organization stops merely counting AI use and starts learning from it because every use case now sits within a structure that can distinguish a low-stakes draft from a decision that requires a human signature before it goes anywhere.
Ubiquitous Access, Real Guardrails
The practical direction that follows from all of this is broad, sanctioned access within real guardrails, not narrower access defended more aggressively.
The safest path should also be the easiest. An employee with a legitimate need should never have to resort to a consumer workaround because the approved service is held up by a slow exception process that assumes every request is exceptional. The organizational environment should provide immediate access for common, low-consequence uses, a clear and fast escalation path for higher-consequence uses, and training that is close enough to the actual work to be used rather than filed away after onboarding.
Guardrails, when done well, protect outcomes, data, authority, and evidence. They do not try to predict every possible prompt or force everyone into a single workflow that fits no one’s actual task. Concrete examples include defined data classes that may or may not be shared with a given tool, required human authorization before a consequential decision is implemented, evidence retention whenever AI materially shapes a result, and verification requirements scaled to the potential harm a wrong answer could cause. Controls like these give people real room to work while keeping the organization accountable for what happens inside that room.
Training matters here too, not the kind built entirely around a list of prohibitions. Access without understanding breeds false confidence, and false confidence is its own risk. People need to understand, concretely, how these tools can produce fluent, confident, and incorrect answers, how they can quietly reproduce bias baked into their training data, and when a claim needs independent verification before anyone acts on it. Training that only tells people what they cannot do will not build the judgment that makes transparent, safe use the obvious choice rather than the inconvenient one.
Disclosure Has to Be a Normal Condition, Not a Confession
None of this works if people are afraid to say what they actually did.
A system cannot learn from AI use that nobody will admit to. If disclosure automatically triggers punishment, people will hide the prompts, the outputs, and the workaround itself, and they will be entirely rational to do so. Technical monitoring can show that a service was accessed. It cannot reconstruct the chain that connects a prompt, a person’s interpretation of the output, their judgment about whether to trust it, and the action they eventually took. That chain is exactly what governance needs to see, and it only becomes visible if the person involved is willing to describe it honestly.
The target condition is assured disclosure. Someone should be able to say that AI contributed to a result without that statement functioning as an admission of guilt. Disclosure should trigger a level of review that matches the actual outcome at stake, not a reflexive search for someone to blame.
After disclosure, the useful questions are what need drove the behavior, what the organizational system enabled or blocked, what evidence actually supports the result, and what should change before the next time. None of that excuses genuinely reckless behavior. It simply separates learning from punishment long enough for the organization to see its real operating system instead of the one it wishes it had.
This is where culture stops being a soft concern and becomes the precondition for governance to work at all. An organization that has previously punished honest disclosure has, correctly, taught its people that disclosure is dangerous. That lesson does not reverse itself just because a new policy says disclosure is now welcome. Assured disclosure has to be demonstrated repeatedly before it will be trusted, and it will not be trusted the first time someone tests it and gets burned.
Not Every Case Is the Same Case
None of this is an argument for treating every instance of unauthorized AI use as an innocent workaround awaiting understanding. Three genuinely different situations are lumped together in the same conversation, and they deserve three distinct responses.
Most of what gets labeled Shadow AI is a legitimate capability gap. Someone needed something the organization did not provide quickly enough, at a usable quality, or with acceptable friction, so AI became the fastest available substitute. The fix is better access, better design, and better training, not a harder line on enforcement.
Some of it is closer to negligence. Access and training may already be adequate, and the person simply exercised poor judgment. No governance system eliminates the possibility of a bad individual decision. What it can do is make the constraints, review requirements, and the consequences of that specific failure clearer, and treat repeated mistakes as a signal to fix a pattern rather than as a reason to write one more policy nobody will read.
A small amount of it is genuinely malicious. Someone using AI to deliberately deceive, steal, sabotage, or bypass a control they fully understand is not seeking a better way to do legitimate work. That is adversarial behavior, and it belongs in security, investigation, and accountability, not in a conversation about better tooling.
Calling all three of these Shadow AI and responding to all three with the same restrictive policy guarantees that the policy fails in the majority case, while doing almost nothing about the minority case it was actually designed for.
The System That Makes This Durable
An organization can accept every argument above and still fail to act on it, because good intentions do not endure without a mechanism that continually tests whether they are true. That mechanism has a name in assurance work. It asks two linked questions about any outcome that matters. What must be true here for this to count as safe and correct? And what is the smallest set of observations that would actually prove that, rather than merely assert it?
Applied to AI-assisted work, this prevents the conversation from drifting back into abstraction. It is not enough to say that the use of AI should be generally safe and authorized. The organization has to specify, for a specific class of work, what evidence would prove that the result was accurate, that a human retained real authority over the outcome, and that the process could be reviewed after the fact rather than trusted on faith.
Within this discipline, AI output is not treated as a decision or as evidence by default. It becomes useful once its origin, verification, and the human judgment applied to it are all visible. That single discipline lets an organization grant broad access without ever confusing access with permission or a fluent answer with a correct one.
This is also why detection alone was never going to be the answer. Detection tells you a service was reached. It tells you nothing about the outcome someone was pursuing, the evidence that shaped the result, or whether they hit a real capability gap the organization still has not closed. Transparency provides the context that detection cannot. An organization that invests only in better detection will get better at counting a problem it still does not understand.
Shining the Light, Not Chasing the Shadow
The public label can stay “Shadow AI” because it gets attention, and it should. But attention is only useful if it points to something true. Shine an actual light on most of what hides under that name, and the monster dissolves into something a leader already knows how to fix. It shows up when people conclude, correctly or not, that the organization will not give them what they need to do legitimate work at the speed the work requires.
The response to that diagnosis is neither unrestricted access nor tighter restrictions. It is broad, usable access within real guardrails, paired with outcome-based assurance that measures the gap between what people asked for and what they received.
Make disclosure a routine condition rather than a confession. Keep human authority explicit and visible in every consequential decision. Preserve the evidence when AI materially shaped a result that matters. Then use that evidence to improve the approved path until going around it is no longer the faster option.
That is a harder discipline than writing a policy and blocking a list of domains. It is also the only version of this that produces an organization capable of learning something true about how its own people are actually working, rather than one that only knows what its paperwork says should be happening.
The shadow was never the thing worth fearing. It was only the absence of light, and light is the one thing an organization has always had the power to turn on itself.
About the Author

Dave is the Executive Director of the DVMS Institute.
Dave spent his “formative years” on US Navy submarines. There, he learned complex systems, functioning in high-performance teams, and what it takes to be an exceptional leader. He took those skills into civilian life and built a successful career leading high-performance teams in software development and information service delivery.
Digital Value Management System® is a registered trademark of the DVMS Institute LLC.
® DVMS Institute 2026 All Rights Reserved


