IndustryAnalysis

Shadow AI: why blocking fails and what actually governs it

Shadow AI now reaches desktop apps, coding agents and local MCP servers that network blocks never see. What the breach data shows, what regulators expect, and the control model that works.

Illustration of a company laptop panel listing a desktop chat app, AI in the browser, a coding agent and a local MCP server, each routed by a blue line into an AI gateway panel labelled Policy with checks for keys, budgets, guardrails and logs; a dashed red personal-account tile leaks out and is stopped by a crossed circle, beside chips reading 1 in 5 and $670K.
Illustration of a company laptop panel listing a desktop chat app, AI in the browser, a coding agent and a local MCP server, each routed by a blue line into an AI gateway panel labelled Policy with checks for keys, budgets, guardrails and logs; a dashed red personal-account tile leaks out and is stopped by a crossed circle, beside chips reading 1 in 5 and $670K.

Shadow AI is the use of AI tools, models, agents and Model Context Protocol (MCP) servers at work without the approval or oversight of the IT and security teams responsible for the data they touch. In IBM’s 2025 Cost of a Data Breach study of 600 organisations, one in five reported a breach due to shadow AI, and only 37% had policies to manage AI or detect it. The surfaces involved have multiplied since the first wave of ChatGPT bans in 2023: desktop chat apps, AI inside browsers, coding agents in the terminal, and local MCP servers that run as subprocesses on a laptop.

This analysis covers what the evidence says about the size of the shadow AI problem, why network blocks and policy memos no longer contain it, what the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001 expect, and the control model that holds up: an AI gateway as the policy engine, extended to every endpoint. Bifrost, an open-source AI gateway written in Go by Maxim AI, is used as the worked example because it now covers both halves of that model.

What is shadow AI, and what it is not

Shadow AI is any AI usage the organisation cannot see or control: an unapproved app, an approved app used through a personal account, a model called with a personal API key, or a tool server wired into an agent without review. It is a subset of shadow IT. The distinction matters because a file-sharing app receives data occasionally, while an AI tool receives data as the input to every interaction.

Three clarifications keep the definition useful.

  • The label describes the channel, not the product. ChatGPT through a company workspace with logging and guardrails is sanctioned. The same model through a personal login on chatgpt.com, on a company laptop, is shadow AI.
  • Malice is not required. The typical case is an employee trying to work faster. Microsoft and LinkedIn’s 2024 Work Trend Index, a survey of 31,000 knowledge workers in 31 markets fielded between 15 February and 28 March 2024, found 75% used generative AI at work and 78% of those users brought their own AI tools. At small and medium-sized companies the figure was 80%.
  • Agents and tools count. A coding agent running on a developer’s own API key, or an MCP server that gives it shell and file access, is shadow AI even though no browser is involved. These are the cases most inventories miss.

How big the shadow AI problem is: what the data shows

The best evidence on shadow AI comes from breach studies and from vendors that observe traffic directly. Each measures something different, so the conditions matter as much as the headline number. The table below lists only reports read in full for this piece.

SourceMethodKey shadow AI findings
IBM Cost of a Data Breach 2025Ponemon study of 600 breached organisations, March 2024 to February 20251 in 5 reported a breach due to shadow AI; high shadow AI added USD 670,000 to breach cost; only 37% had policies to manage or detect it
Netskope Cloud and Threat Report 2026Anonymised usage data from the Netskope One platform47% of genAI users use personal AI apps (down from 78%); 223 genAI data policy violations per organisation per month
LayerX Enterprise AI and SaaS Data Security Report 2025Browser telemetry from the vendor’s enterprise customers67% of genAI access through non-corporate accounts; 82% of data pasted into genAI comes from unmanaged accounts
Microsoft and LinkedIn Work Trend Index 2024Survey of 31,000 knowledge workers, 31 markets78% of AI users bring their own AI tools to work

The breach data

IBM’s figures are the most cited and the most carefully conditioned. The USD 670,000 is the difference in average breach cost between organisations with high levels of shadow AI and those with low levels or none, against a global average breach cost of USD 4.44 million and a US average of USD 10.22 million. Separately, 13% of organisations reported breaches of AI models or applications, and of those, 97% said they had no AI access controls in place. That 97% applies to the breached subset, not to all organisations, and is often misquoted.

The governance numbers are the more telling ones. IBM reports that 63% of breached organisations either had no AI governance policy or were still developing one, and that among those with policies, only 34% performed regular audits for unsanctioned AI. A written policy without an audit is a statement of intent.

The traffic data

Netskope’s 2026 report shows the shape of the problem changing rather than shrinking. The share of AI users on personal AI apps fell from 78% to 47% as enterprises rolled out sanctioned tools, but the share switching between personal and enterprise accounts rose from 4% to 9%. Over the same year the number of genAI users in the average organisation tripled and prompts grew sixfold, from 3,000 to 18,000 a month. Data policy violations doubled: 3% of genAI users generated an average of 223 incidents a month. Source code made up 42% of those violations, regulated data 32% and intellectual property 16%.

LayerX’s browser telemetry points at the mechanism. It reports that 77% of employees paste data into generative AI tools, that 82% of that pasting comes from unmanaged accounts, and that 22% of those employees paste data containing personal or payment card information. Copy and paste leaves no file for file-centric data loss prevention to inspect. Both Netskope and LayerX sell controls for this problem, and their figures come from their own customer bases, not a random sample.

What the numbers do not show

None of these sources measures coding agents or local MCP servers directly. Netskope names AI browsers and MCP-enabled agents as emerging risks for 2026 but does not quantify them. The measured problem is the browser and SaaS portion; the agentic portion is growing in a blind spot most measurement tools share.

Shadow AI risks: what leaks, and why agents change the picture

The main shadow AI risks fall into five groups, and the agentic ones carry the largest blast radius.

  1. Data exposure. Prompts carry source code, customer records and internal documents to providers under consumer terms, outside the organisation’s contracts and retention settings.
  2. Credential leakage. API keys and tokens pasted into prompts, or read from .env files by an agent, leave the machine in plain text.
  3. Uncontrolled spend. Personal API keys expensed monthly produce no per-team budget and no usage attribution.
  4. Unvetted tool execution. MCP servers can read files, call APIs and run commands on the user’s behalf.
  5. Compliance gaps. An organisation cannot document, assess or audit an AI system it does not know exists.

The fourth group deserves attention because it is new. The MCP transport specification defines stdio as a standard transport in which “the client launches the MCP server as a subprocess,” and says clients should support stdio whenever possible. The protocol’s own security best practices warn that local MCP servers “may have direct access to the user’s system,” list arbitrary code execution and data exfiltration among the risks, and give the example of a malicious startup command embedded in a client configuration. A developer who adds an MCP server to Claude Code or Cursor has installed software that runs with their privileges and talks to the agent over a pipe, not a network. The explainer on how an MCP gateway works covers the server-side half of this problem.

The earliest well-documented incident shows how quickly this escalates. In May 2023, after engineers uploaded sensitive code to ChatGPT, Samsung restricted generative AI on company-owned computers, tablets and phones, and on non-company devices using internal networks. An internal survey found about 65% of participants thought generative AI tools carried a security risk. The ban was described as temporary, pending secure internal tooling. Incident, ban, sanctioned alternative: most enterprises have followed that sequence since, and the ban stage is where shadow AI grows.

Why network blocking and policy memos fail

Network blocking and acceptable-use memos fail because they assume AI traffic crosses a point the organisation controls and that people remember the rules under deadline pressure. Neither assumption holds for current AI tooling. The figure below maps the five channels by which shadow AI leaves a company machine.

Five columns of cards: browser AI via personal logins and copy-paste, desktop apps installed per user, coding agents with personal API keys, local MCP servers launched as subprocesses with no network hop, and AI features added inside approved SaaS tools

Figure 1: Each channel defeats a different control, and the two agentic channels are invisible to network inspection altogether.

The perimeter is not where the laptop is

A DNS or proxy deny-list only applies when the laptop is on the corporate network or tunnelled through a secure web gateway. Remote and hybrid work means many AI requests leave from home networks. Where traffic is inspected, the domain is often the same for sanctioned and unsanctioned use: a personal and an enterprise session on the same chat service differ by account, not destination. Netskope’s finding that 9% of users switch between personal and enterprise accounts is exactly the case a domain block cannot separate.

Blocking moves usage rather than stopping it

Netskope reports that 90% of organisations actively block at least one genAI app, the average blocking 10. That suits tools with no business purpose. As a primary control it fails predictably: the user moves to a tool not yet on the list, or to a personal phone.

Local traffic never touches the network

A stdio MCP server exchanges messages with its client over standard input and output. There is nothing for a proxy to see until the server itself makes an outbound call, and a filesystem or shell server may never do so. Coding agents add a further gap: the model call goes to a provider API with a key set in the developer’s shell, which looks like any other HTTPS client.

Memos have no telemetry

An acceptable-use policy is necessary because it defines what “approved” means, but it produces no inventory, no logs and no enforcement. IBM’s 34% audit rate is the measure of how rarely policy is checked against reality.

Three rows comparing responses: a policy memo relies on recall and is unmeasured; network blocking covers known domains only and pushes use elsewhere; a gateway plus endpoint agent routes at the machine and leaves traffic governed and logged

Figure 2: The first two responses stop where the corporate network stops; routing at the machine does not.

What regulators and standards expect

Regulators and standards bodies do not use the phrase “shadow AI,” but each framework assumes the organisation knows which AI systems its people use. The table maps the provisions most relevant to unsanctioned use, as of September 2026.

FrameworkProvisionWhat it expectsWhy shadow AI breaks it
EU AI ActArticle 4 (as amended July 2026)Providers and deployers take measures to support AI literacy of staffCannot target training at tools nobody has disclosed
EU AI ActArticle 26 (high-risk deployers)Use per instructions, competent human oversight, logs keptUnknown systems cannot be overseen or logged
NIST AI RMF 1.0GOVERN 1.6, GOVERN 6.1Inventory of AI systems; policies for third-party AI riskNo inventory exists for unsanctioned tools
ISO/IEC 42001:2023Annex A.4, A.9, A.10Documented AI resources and tooling; responsible-use processes; supplier responsibilitiesTools and suppliers outside the management system

The EU AI Act after the Digital Omnibus

The original Article 4 of Regulation (EU) 2024/1689, in force since 2 February 2025, required providers and deployers to ensure “to their best extent, a sufficient level of AI literacy” among staff. The Digital Omnibus, Regulation (EU) 2026/1744, in force since 27 July 2026, rewrote it. As summarised by the Future of Privacy Forum and Gibson Dunn, providers and deployers now “shall take measures to support the development of AI literacy,” and the text clarifies that this does not require guaranteeing any specific level for any individual. The duty became one of effort rather than result, but it still sits on every deployer. Taking proportionate measures presupposes knowing what tools staff use.

For high-risk systems, Article 26 places concrete duties on deployers, including assigning human oversight to people with the necessary competence and keeping automatically generated logs. The Omnibus moved the Annex III high-risk dates to 2 December 2027, so these duties are not yet live for most systems. The one-year review of the Act’s model rules sets out the full timeline. The practical point is that a team preparing for 2027 needs an inventory now, and shadow AI is by definition missing from it.

NIST AI RMF and ISO/IEC 42001

The NIST AI RMF 1.0 is voluntary but widely used as a reference. GOVERN 1.6 reads: “Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities.” GOVERN 6.1 asks for policies addressing risks from third-party entities, and GOVERN 2.2 for AI risk management training for personnel and partners. A personal chatbot account is a third-party AI system outside the inventory.

ISO/IEC 42001:2023, the certifiable AI management system standard, groups its Annex A controls under nine objectives. Three bear on shadow AI: A.4 (resources for AI systems, including A.4.4 tooling resources), A.9 (use of AI systems, including A.9.2 processes for responsible use) and A.10 (third-party and customer relationships). An auditor testing A.9 will ask how the organisation knows its responsible-use processes are followed. Logs from a governed path are the most direct answer.

The control model: an AI gateway as the policy engine

An AI gateway is a service between applications and model providers that authenticates callers, applies policy and records each request. For shadow AI, it is where “approved” becomes enforceable: every sanctioned AI request carries an identity, draws on a budget, passes guardrails and leaves a log entry. The definitional explainer on what an AI gateway is covers the general architecture; this section covers the controls that matter for governance.

Bifrost is the strongest fit for this role among the gateways assessed here, for reasons that come down to verifiable capabilities rather than branding. Its governance model is built on virtual keys, which carry model and provider filtering, independent budgets, rate limits and an on/off switch, so access can be revoked without touching provider credentials. The same keys hold a strict allow-list of MCP tools through per-key MCP tool filtering, which is deny-by-default: a key with no MCP configuration gets no tools except those explicitly marked as allowed by default.

The security layer sits in the same request path. Bifrost guardrails validate inputs and outputs for model traffic and MCP tool executions, with native Gitleaks-backed secrets detection, custom regex rules including a PII template, and integrations such as AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor and CrowdStrike AIDR. Every request and response is captured by built-in observability with tokens, cost and latency, while signed audit logs record administrative changes such as who created a key or edited a guardrail. That split, request logs for usage and audit logs for policy changes, is what an ISO/IEC 42001 or EU AI Act review will ask for.

Bifrost is also open source and self-hostable, so prompts and logs can stay inside the organisation’s own network; the survey of open-source LLM gateways you can self-host compares deployment and licensing options across the category. For a broader view of routing, failover and overhead, the comparison of production LLM gateways and GenAI Brief’s own five-gateway comparison set Bifrost against the alternatives.

The limit of any gateway, Bifrost included, is structural. It governs only the traffic configured to flow through it. An application team that sets a base URL is covered. An employee who installs a desktop chat app, or a developer whose coding agent reads a personal key from the shell, is not. That is the gap shadow AI occupies.

Extending the gateway to the endpoint

Endpoint enforcement closes the gap by moving the routing decision from each application’s configuration to the machine itself. An agent on each company laptop sends AI traffic from supported apps through the gateway, whatever network the laptop is on, and blocks apps and MCP servers the organisation has not approved. The gateway stays the policy engine; the endpoint agent makes sure the traffic reaches it.

Bifrost Edge implements this model as the endpoint layer of Bifrost. It is in alpha as of September 2026, with teams registering for onboarding, so it should be evaluated as early-access software rather than a finished product. The design is what makes it relevant here: Edge enforces the virtual keys, budgets, guardrails and logging already configured in the gateway, rather than introducing a second policy system.

Pipeline from admin policy through MDM push, SSO sign-in and an app on the device to a gateway check and the provider with logging, with branches showing a denied app stopped on the device and a secret or PII blocked or redacted at the gateway

Figure 3: Policy is authored once in the gateway; the endpoint agent’s job is to make every supported AI request pass through it.

What Edge covers today

Edge runs on macOS, Windows and Linux. After a one-time browser sign-in through the organisation’s single sign-on, it sits in the menu bar or system tray and routes AI traffic at the machine level, with no base URLs to change. Its supported applications list includes Claude Desktop, the ChatGPT and Codex desktop apps and Cursor; Claude Code, Codex CLI and OpenCode as coding agents; and ChatGPT and Claude on the web. The list is growing, and anything not on it is outside Edge’s routing today.

App and MCP governance

App governance lets administrators allow or block AI applications centrally. Allowed apps run normally with traffic governed through Bifrost; disallowed apps are blocked before data leaves the machine. When Edge detects a new app or MCP server, it requests approval in the admin console, and administrators choose whether pending items are allowed or blocked meanwhile. That setting is the practical dial between disruption and control during rollout.

MCP governance addresses the channel that network tools miss. Edge reads the MCP configuration of supported apps on each machine, including Claude Code, Claude Desktop, Gemini CLI, OpenCode, Codex and Cursor, and builds a fleet-wide inventory of which servers are configured and on how many devices. The catalogue is deduplicated, so a server on 300 laptops is approved or denied once. A denied server is blocked on the device, even in an app that had it configured before the policy existed. For GOVERN 1.6 and ISO/IEC 42001 A.4.4, that inventory is itself evidence.

Guardrails and rollout

Because Edge routes traffic through the gateway, endpoint security needs no separate configuration: the same guardrail profiles apply to a prompt typed into ChatGPT in the browser as to an internal application. Secrets and PII are caught before the prompt reaches a provider.

Rollout is designed around existing device management. MDM deployment supports Jamf, Microsoft Intune, Kandji, Omnissa Workspace ONE and JumpCloud. The managed configuration carries only the gateway and management endpoints; identity and keys come from the user’s SSO sign-in. Edge also needs an organisation certificate on each machine, because it routes encrypted AI traffic through the gateway. Security teams should review that certificate’s handling as they would any TLS inspection deployment.

Limits of the endpoint model

Three limits apply. Edge governs managed company machines, so personal phones and home computers remain a policy and training question. Coverage depends on the supported-app list, so an obscure desktop client may pass unrouted until it is added. And as alpha software, Edge should be piloted with a volunteer group before any fleet-wide deny policy is switched on.

AI governance tools compared by layer

AI governance tools fall into layers, and most organisations need more than one. The table compares them on the question that matters for shadow AI: which channels each layer can see and enforce.

LayerSees browser AISees desktop appsSees coding agentsSees local MCPEnforces budgets and guardrails
Acceptable-use policy and trainingNoNoNoNoNo
Secure web gateway or CASBOn inspected networksPartly, by domainPartly, by domainNoData rules only
Browser security extensionYesNoNoNoData rules only
AI gateway aloneConfigured apps onlyConfigured apps onlyIf pointed at itServer-side MCP onlyYes
AI gateway plus endpoint agentSupported sitesSupported appsSupported agentsDiscovered and enforcedYes

Network and browser tools see data but not identity, budgets or tool use; a gateway alone sees everything in its path but cannot reach onto the laptop. The combination is the only row with a yes in the last column and coverage across all four agentic channels, which is why the better choice for most engineering-heavy organisations is to build governance around a gateway and extend it to the endpoint. Teams comparing gateway candidates for that core can use the list of leading AI gateways for production as a starting shortlist.

How to prevent shadow AI: a 90-day plan

Preventing shadow AI means making the governed path the easiest one and then measuring the rest. The sequence below suits an organisation of a few hundred to a few thousand managed machines.

  1. Weeks 1 to 2: publish the approved list. A one-page acceptable-use policy naming sanctioned tools, prohibited data classes and how to request a new tool. Short enough to remember.
  2. Weeks 2 to 4: stand up the gateway. Deploy Bifrost in the organisation’s own cloud or data centre (the self-hosted AI gateway options guide covers Kubernetes and air-gapped setups), connect providers, issue virtual keys per team and turn on secrets and PII guardrails.
  3. Weeks 4 to 6: move sanctioned traffic. Point internal applications and approved coding agents at the gateway. Set team budgets from a month of observed spend rather than guesses.
  4. Weeks 6 to 8: pilot endpoint enforcement. Deploy the endpoint agent through MDM to a volunteer group, with pending apps and MCP servers allowed. Use the discovered inventory to find what people actually use.
  5. Weeks 8 to 10: decide the catalogue. Approve the tools with a business case, deny the rest, and add the approved ones to the policy. Bulk-deny pending MCP servers only after owners have had a chance to request approval.
  6. Weeks 10 to 13: expand and audit. Roll out fleet-wide, switch pending items to blocked, and schedule a quarterly review of logs, denied requests and budget use. That review is the “regular audit for unsanctioned AI” that only 34% of IBM’s policy-holding organisations perform.

Two scenarios show how the plan lands. A 400-person software company with most staff remote will find that network controls saw little of its AI traffic; the endpoint inventory in step 4 typically becomes the first complete list of coding agents and MCP servers it has ever had. A regulated financial firm that already runs a secure web gateway will get most value from steps 2 and 3, because per-team keys and request logs give it the attribution its model risk function needs, and from MCP governance, which its network tooling cannot provide.

The judgement

Shadow AI is no longer a question of whether employees use ChatGPT. It is a question of whether an organisation can name every AI system touching its data, including the agents and tool servers that run on laptops without crossing a network boundary. The breach data says most cannot, and the frameworks that auditors and regulators will use in 2027 assume they can.

Bans and memos buy time. The durable answer is structural: an AI gateway where access, budgets, guardrails and logs are defined once, and an endpoint layer that routes each machine’s AI traffic through it. Bifrost is the most complete implementation of that model assessed here, with a mature gateway and an endpoint agent that is still in alpha and should be piloted accordingly. Teams working through shadow AI governance can request a Bifrost demo or start with the open-source gateway and add endpoint enforcement once the sanctioned path is in place.

Feature claims are drawn from each vendor’s public documentation as of September 2026.

Sources

  1. IBM Report: 13% of organizations reported breaches of AI models or applications, 97% of which reported lacking proper AI access controls IBM Newsroom
  2. 2025 Cost of a Data Breach Report: Navigating the AI rush without sidelining security IBM
  3. 2024 Work Trend Index: AI at work is here. Now comes the hard part Microsoft and LinkedIn
  4. Cloud and Threat Report: 2026 Netskope Threat Labs
  5. Enterprise AI and SaaS Data Security Report 2025 LayerX
  6. Samsung bans use of generative AI tools like ChatGPT after April internal data leak TechCrunch
  7. Model Context Protocol specification: Transports Model Context Protocol
  8. Model Context Protocol: Security Best Practices Model Context Protocol
  9. Regulation (EU) 2024/1689 (the AI Act) EUR-Lex
  10. Regulation (EU) 2026/1744 amending the AI Act (the Digital Omnibus on AI) EUR-Lex
  11. The AI Act implementation timeline: what changes under the AI Omnibus? Future of Privacy Forum
  12. EU AI Act Omnibus agreement: postponed high-risk deadlines and other key changes Gibson Dunn
  13. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 National Institute of Standards and Technology
  14. ISO/IEC 42001 Annex A controls explained ISMS.online
  15. Bifrost Edge overview Bifrost documentation

Frequently asked questions

What is shadow AI?

Shadow AI is the use of AI applications, models, agents or MCP servers inside an organisation without the knowledge, approval or oversight of IT and security teams. Typical examples are a personal ChatGPT or Claude account used for work, a coding agent running on a personal API key, or an MCP server an engineer installed locally. It is a subset of shadow IT with higher data risk, because every prompt carries data.

Is ChatGPT considered shadow AI?

ChatGPT is shadow AI when it is used for work through an account or channel the organisation has not approved and cannot see, such as a personal login on chatgpt.com. The same product used through a sanctioned enterprise workspace, or routed through a governed gateway with logging and guardrails, is not shadow AI. The label describes the governance status of the usage, not the tool.

What are the main shadow AI risks?

The main shadow AI risks are data exposure (source code, regulated personal data and intellectual property sent to providers under consumer terms), credential leakage through pasted secrets, uncontrolled spend on personal API keys, unvetted MCP servers executing code with the user's privileges, and compliance gaps, because an organisation cannot document, assess or audit AI systems it does not know exist.

How do you detect shadow AI?

Detection starts with an inventory: network and SaaS logs show known AI domains, expense reports show personal subscriptions, and endpoint tooling shows installed AI apps and configured MCP servers. Network data alone misses off-network laptops and local MCP servers, so an agent on each managed device that reports apps and MCP configurations gives the most complete picture.

How can organisations prevent shadow AI?

Organisations prevent shadow AI by making the governed path the easiest one: provide approved AI tools, route them through an AI gateway that applies keys, budgets, guardrails and logging, then deploy an endpoint agent through MDM so every AI app on company machines uses that gateway and unapproved apps or MCP servers are blocked. Training and a short acceptable-use policy support this but do not replace it.

What is the difference between shadow IT and shadow AI?

Shadow IT is any technology used without IT approval, such as an unsanctioned file-sharing app. Shadow AI is the AI-specific subset. The difference in practice is that AI tools consume data as their input on every interaction, and agents and MCP servers can take actions, so the exposure is continuous rather than occasional and the blast radius includes systems the tool can act on.

What are AI governance tools?

AI governance tools are software that enforce and evidence policy over how AI is used: inventories of AI systems, AI gateways that control access, budgets and guardrails for model traffic, endpoint agents that bring desktop and agent traffic under those controls, and audit logging that supports frameworks such as the EU AI Act, the NIST AI RMF and ISO/IEC 42001.

All industry →