ToolsExplainer

MCP gateway explained: what it does and when agent teams need one

An MCP gateway puts one governed endpoint between agents and the MCP servers they call. How tool filtering, upstream auth, audit and token control work, and where the risks sit.

Illustration of a glass policy card titled 'Who may call what' listing six MCP tool calls such as github.create_issue and shell.exec, each marked ALLOW in green or DENY in red, beside a tilted blue virtual-key card reading '4 of 38 tools' and a chip reading '−92.8% input tokens, 508 tools'.
Illustration of a glass policy card titled 'Who may call what' listing six MCP tool calls such as github.create_issue and shell.exec, each marked ALLOW in green or DENY in red, beside a tilted blue virtual-key card reading '4 of 38 tools' and a chip reading '−92.8% input tokens, 508 tools'.

An MCP gateway is a proxy that sits between AI agents and the Model Context Protocol servers they call, so that tool discovery, credentials, access policy and logging are handled in one place instead of being configured separately inside every agent. It matters once more than a handful of agents share more than a handful of MCP servers: at that point each agent host holds its own tokens, sees every tool whether it needs it or not, and leaves no shared record of what it did. This explainer is for platform and security engineers deciding whether to put one in front of their tools, and what it should do when they do.

The worked example throughout is Bifrost, an open-source AI gateway written in Go by Maxim AI that acts as both an MCP client to upstream servers and an MCP server to agents, alongside routing ordinary LLM traffic. The mechanics described here are general; the specifics are checked against each project’s documentation and the MCP specification as of September 2026.

What is an MCP gateway?

An MCP gateway is an aggregation and policy layer for MCP tools. Agents connect to a single MCP endpoint on the gateway. The gateway connects to the real MCP servers, merges their tools into one catalogue, and on each request decides which of those tools the caller can list and call, which credential to use upstream, and what to record.

Three neighbouring pieces of infrastructure are easy to confuse with it, and the differences decide where each one belongs.

LayerWhat it governsTypical controlsExample question it answers
MCP serverExecution of its own toolsInput validation, its own auth“Can this server create a Jira ticket?”
MCP gatewayWhich agent reaches which tool on which serverTool allow-lists, upstream credentials, tool-call logs, argument checks“May the support agent call shell.exec?”
LLM gatewayTraffic to model providersProvider keys, budgets, failover, prompt guardrails“Which model serves this request, and who pays?”
API gatewayHTTP traffic to known servicesRoutes, rate limits, API keys“Is this client allowed to hit /orders?”

In practice the MCP and LLM gateway roles are increasingly shipped together, because the same identity (a key, a team, a user) should govern both the model calls and the tool calls an agent makes. The plain-language guide to AI gateways covers the model side; the comparison of production LLM gateways shows which products include native MCP support and which bolt it on through plugins.

Why agent teams add an MCP gateway

Teams add an MCP gateway when the number of connections, not the number of tools, becomes the problem. With N agents and M servers wired directly, there are up to N×M configurations, each with its own credential, its own copy of the tool list and its own logging, or none. The protocol itself does not centralise any of that.

The specification is explicit about why. The current revision (2025-11-25) defines two standard transports: stdio, where the client launches the server as a local subprocess, and Streamable HTTP, which replaced the older HTTP+SSE transport from the 2024-11-05 revision. Authorisation is marked OPTIONAL. HTTP servers that implement it should follow an OAuth 2.1 profile, while stdio implementations “SHOULD NOT follow this specification, and instead retrieve credentials from the environment.” In a point-to-point estate, that means API tokens sitting in configuration files and environment variables on every machine that runs an agent.

The consequences show up in four places:

  • Credential sprawl. A GitHub token with repository write access lives in every developer’s IDE config and every CI agent’s environment. Rotating it means finding all of them.
  • Tool sprawl. Each agent sees every tool on every server it is connected to, including ones it should never call.
  • No shared record. The spec says clients SHOULD “log tool usage for audit purposes”, but each client logs locally, in its own format, if at all.
  • No single place to change policy. Revoking one team’s database access means editing every agent it runs.

Two rows compared. Top row, without a gateway: each agent wired to each server, tokens held on hosts in config files, every tool in the model context on every turn, and no shared record of who called what. Bottom row, with a gateway: one endpoint scoped per key, credentials held centrally, tools loaded on demand through a filtered list or Code Mode, and every call logged plus an admin audit trail.

Figure 1: The agents and servers do not change. A gateway moves credentials, tool lists and records from every host to one place.

A gateway also changes who can reason about an agent’s reach. In the agent frameworks survey, MCP is the shared tool protocol across all four framework camps; a gateway is what lets a platform team set tool policy once regardless of which framework each product team picked.

How an MCP gateway handles a tool call

An MCP gateway processes a tool call in six steps: identify the caller, filter the tool list, check the specific call, attach the upstream credential, forward it, and log the result. Each step is a place where policy applies, and each is where products differ most.

Pipeline from an agent or IDE connecting to one MCP URL with one key, through identify caller, filter tools, check the call, the MCP server with the gateway holding the credential, and log the call. Dashed branches show an expired key receiving a 403 with no tools, and a denied call being stopped before execution.

Figure 2: Identity, filtering and argument checks all run before an upstream server sees anything.

Tool discovery and filtering

The first job is aggregation. Bifrost as an MCP gateway connects to upstream servers over stdio, HTTP or SSE and exposes the combined catalogue at a single /mcp endpoint: POST for JSON-RPC messages such as tools/list and tools/call, GET for a server-sent-events stream. It pings each upstream server every 10 seconds by default with a 5-second timeout, marks a server unstable after five consecutive failures, and re-syncs each server’s tool list on a default 10-minute cycle so newly added tools appear without a restart.

Filtering then stacks in layers. Bifrost applies three levels of tool filtering: the server configuration sets a baseline (tools_to_execute, where an empty list means no tools), a request header can narrow it for one call, and the virtual key narrows it again. A tool must pass every applicable layer to reach the model.

Per-key tool allow-lists

The control that matters most for security is the per-caller allow-list. In Bifrost, a virtual key with no MCP configuration gets no MCP tools: deny by default, except for servers an administrator has explicitly marked “Allow by Default”. For each server attached to a key, the administrator lists specific tools or uses * for all. The tools/list response contains only what the key allows, tools/call is checked against the same list at execution time, and an inactive or expired key is refused with HTTP 403. A caller can send an x-bf-mcp-include-tools header to narrow its own list further, but the header “can only narrow, never widen” the key’s grant.

For teams that want a curated bundle rather than a per-key list, Bifrost’s Virtual MCPs group selected tools from several servers behind their own endpoint at /mcp/<slug>, reachable only through the virtual keys they are attached to. The feature, previously called MCP tool groups, is part of the open-source gateway; the enterprise tier adds grants through access profiles and role-based visibility.

For example, a support agent’s key is attached to a Virtual MCP containing jira.search, github.create_issue and a read-only Postgres tool. The same gateway also fronts filesystem and shell servers, but the support agent’s tools/list never mentions them, and a crafted tools/call for shell.exec fails the allow-list check before any subprocess starts.

Authentication to upstream servers

The second security job is credential custody. The specification forbids “token passthrough”: an MCP server “MUST NOT pass through the token it received from the MCP client” to an upstream API, and must reject tokens that were not issued for it. The security best practices spell out why: passthrough breaks audience validation, hides the real caller from downstream logs and turns the server into an exfiltration proxy for anyone holding a stolen token.

A gateway resolves this by holding upstream credentials itself and deciding whose identity to use. Bifrost supports six MCP authentication types:

Auth typeWho authenticatesSuits
noneNobodyPublic or local tools with no key
headersAdministrator, onceA shared internal service token
oauthAdministrator, onceOne company-wide OAuth app
per_user_headersEach user, on first callPersonal API keys
per_user_oauthEach user, on first callPersonal Notion, GitHub or Sentry accounts
token_exchange (enterprise)Each caller, every callInternal servers that trust the company identity provider

The per-user modes matter for least privilege: a shared admin token gives every agent the union of everyone’s access, while per-user OAuth means the agent acting for a support engineer can only see that engineer’s repositories. Bifrost stores per-user credentials against the caller’s identity and, with token exchange, stores none at all; offboarding a user at the identity provider takes effect within the cached token lifetime, which the documentation caps at five minutes.

Audit of tool calls

The last job is the record. The OWASP MCP Top 10 lists “Lack of Audit and Telemetry” as MCP08, and the reason is practical: when an agent does something unexpected, the investigation starts with which tool it called, with which arguments, under whose identity.

Two kinds of log answer different questions. Tool-call logs record each execution; Bifrost writes MCP log entries alongside LLM logs and can attach selected request headers as metadata for tracing and tenant identification. Administrative audit logs record changes to policy itself, such as who widened a key’s tool list; in Bifrost Enterprise those entries can be HMAC-signed, retained on a schedule and exported as JSON, JSON Lines or Syslog. A team that only has the first kind can see what an agent did but not who gave it permission.

One default inverts what many agent SDKs do. Bifrost does not execute tool calls automatically: a chat completion returns suggested calls, and execution needs a separate API call unless Agent Mode is switched on for named tools. In pure gateway mode, approval stays in the host application, which matches the specification’s guidance that there “SHOULD always be a human in the loop with the ability to deny tool invocations.”

The token cost of large tool lists

Every tool definition an agent can see is sent to the model on every turn, so tool count becomes a line item. A gateway can cut it in two ways: send fewer definitions through filtering, or replace direct tool calls with a code-execution pattern that loads definitions only when needed.

Anthropic described the problem in its November 2025 post on code execution with MCP: tool descriptions “occupy more context window space, increasing response time and costs,” and “every intermediate result must pass through the model.” In its example, presenting tools as code files so the agent loads only the definitions it needs cut token usage from 150,000 to 2,000, a saving of 98.7%. That is a single illustrative case, not a benchmark.

Bifrost implements the pattern as Code Mode. Instead of the full catalogue, the model sees four meta-tools (listToolFiles, readToolFile, getToolDocs and executeToolCode), reads Python-style stubs for the servers it needs, and writes a short Starlark script that Bifrost runs in a sandbox with the real tool bindings. Intermediate results stay in the sandbox and only the final output returns to the model. Code Mode is enabled per upstream server, so small utilities can stay as direct tools.

The vendor’s published benchmark, run on Claude Sonnet 4.6 with the same query set in each round, is the most detailed public data on the effect:

RoundMCP footprintPass rate, classicPass rate, Code ModeInput tokens, classicInput tokens, Code ModeChange
196 tools, 6 servers64/6464/6419.9M8.3M−58.2%
2251 tools, 11 servers64/6565/6535.7M5.5M−84.5%
3508 tools, 16 servers65/6565/6575.1M5.4M−92.8%

Estimated cost fell by 55.7%, 83.4% and 92.2% across the three rounds. Two caveats apply. First, this is a vendor benchmark, not an independent measurement; the raw report is public in the bifrost-benchmarking repository, which helps. Second, the wall-clock gains in that report are smaller than the token gains: 26.8%, 6.8% and 17.1% faster across the rounds, and in round 1 the simplest queries were marginally slower. The savings scale with tool count, which is why the documentation recommends Code Mode from three or more servers and classic tool calls for one or two.

Code execution also has a security cost that Anthropic states plainly: “Running agent-generated code requires a secure execution environment with appropriate sandboxing, resource limits, and monitoring.” Filtering has no such cost. For many teams the first and cheapest token saving is simply not showing a support agent the 60 tools of an infrastructure server it will never call, and the Bifrost MCP gateway resource page sets out how allow-lists and Code Mode combine.

MCP server security: the risks a gateway addresses

MCP server security has two halves: attacks that arrive through tools, and damage that follows from agents having more access than they need. A gateway does little about the first on its own and a great deal about the second.

The attack that made MCP security a topic is tool poisoning. In April 2025 Invariant Labs demonstrated that instructions hidden in a tool’s description, “invisible to users but visible to AI models,” could make an agent read ~/.cursor/mcp.json and ~/.ssh/id_rsa and send their contents out through an innocent-looking parameter of an add tool. The same post described two variants: a rug pull, where a server changes a tool’s description after the user approved it, and shadowing, where one malicious server’s descriptions alter how the agent uses tools from other, trusted servers. OWASP now tracks the class as MCP03:2025 Tool Poisoning.

The second pattern arrives through tool results rather than descriptions. In May 2025 Invariant Labs showed that a malicious issue filed in a public repository could steer an agent using the GitHub MCP server into reading the user’s private repositories and leaking their contents in a pull request on the public one. The researchers called it a “toxic agent flow”, and their first recommended mitigation was a permission boundary: an agent may access only one repository per session.

That recommendation is the general rule. OWASP’s LLM06:2025 Excessive Agency traces agent damage to three root causes: excessive functionality, excessive permissions and excessive autonomy. Its first mitigation is to “limit the extensions that LLM agents are allowed to call to only the minimum necessary.” Per-key allow-lists, per-user credentials and a human approval step map directly onto those three causes.

The OWASP MCP Top 10, in beta as of September 2026, gives a checklist to map controls against:

OWASP MCP riskWhat a gateway contributes
MCP01 Token Mismanagement and Secret ExposureUpstream credentials held centrally, not on each host
MCP02 Privilege Escalation via Scope CreepPer-key tool allow-lists; per-user auth modes
MCP03 Tool PoisoningOnly reviewed servers reach the catalogue; curated tool sets
MCP06 Prompt Injection via Contextual PayloadsGuardrails on tool arguments and results
MCP07 Insufficient Authentication and AuthorizationOne authenticated endpoint; expired keys refused
MCP08 Lack of Audit and TelemetryTool-call logs plus policy-change audit logs
MCP09 Shadow MCP ServersNothing, unless paired with an endpoint agent

For the injection rows, Bifrost Enterprise applies guardrails at the tool-execution boundary: a rule can inspect or redact a tool’s arguments before it runs and block it before execution, or inspect a successful result and withhold it from the model. Rules target specific servers, tools and even individual arguments, and use the same guardrail providers as LLM traffic.

Four columns. Tool poisoning, MCP03: hidden instructions; a gateway contributes an allow-list and curated set; descriptions still need review. Injection via results, MCP06: payload in tool output; a gateway contributes output guardrails; human approval still needed. Over-broad credentials, MCP01 and MCP02: tokens and scopes; a gateway contributes per-user auth and per-key tools; upstream scopes still need narrowing. Shadow servers, MCP09: unapproved local servers; a gateway alone is blind to laptops; an endpoint agent can inventory and deny them.

Figure 3: A gateway reduces how far an attack can reach. It does not detect a poisoned description on its own.

What a gateway cannot fix

The counter-argument deserves weight. A gateway sees tool names, arguments and results; it does not understand intent. If an approved server ships a poisoned description, the gateway forwards it faithfully unless someone reviewed the text or a guardrail flags it. Allow-lists also have a maintenance cost: they are only as good as the review behind them, and a team that sets every key to * has bought a logging proxy, not a policy layer. And a gateway adds a hop, which matters for chatty agents making dozens of calls per task.

There is also a structural limit: a gateway only governs traffic configured to reach it. The specification’s security guidance describes local MCP servers as binaries that “may have direct access to the user’s system”, started by a client configuration that can embed arbitrary commands. None of that passes through a central gateway.

Laptops and endpoint MCP servers

Developer machines are where most MCP servers actually run. Claude Code, Cursor, Codex and Claude Desktop all let a user add a stdio server with a short configuration entry, and that server runs with the user’s privileges and credentials from the local environment. OWASP lists the result as MCP09, Shadow MCP Servers: tools in use that no one approved and no gateway sees.

This is where the gateway model needs an endpoint component. Bifrost applies governance and security controls (virtual keys, budgets, guardrails, audit logs) centrally at the gateway, and Bifrost Edge, currently in alpha, extends that same governance and security to AI traffic on employee machines. For MCP specifically, Edge MCP governance reads the MCP configuration inside supported AI apps on each machine and builds a fleet-wide inventory: which servers are configured, in which apps, on how many devices. Discovery covers Claude Code, Claude Desktop, Gemini CLI, OpenCode, Codex and Cursor today.

Administrators then allow or deny each server once for the whole fleet. The decision is enforced on the device rather than advised: a denied server cannot be used even by an app that had it configured before the policy existed. Newly discovered servers raise an approval request in the admin console, and administrators choose whether pending servers keep working or are blocked until reviewed. The Edge agent is pushed silently through MDM tools such as Jamf, Intune and Kandji, with Workspace ONE and JumpCloud also documented.

The gateway is the control plane where tool policy, credentials and logs live for agents that connect to it; the endpoint layer finds and governs the servers that were never pointed at it. The wider problem, unsanctioned AI apps and not only MCP servers, is covered in shadow AI governance.

Open-source MCP gateway options in 2026

Open-source MCP gateways split into two groups: gateways that govern both LLM and MCP traffic under one identity model, and dedicated MCP proxies focused on hosting, routing or federating tool servers.

GatewayScopeTool access controlDeployment focusLicence
BifrostLLM and MCP gatewayPer-key allow-lists, deny by default; Virtual MCPs; six upstream auth modes; Code ModeSelf-hosted binary, Docker, KubernetesApache 2.0
Docker MCP GatewayMCP onlyRuns each server in an isolated container with restricted privileges; injects credentialsLocal with Docker Desktop, or Docker EngineOpen source
Microsoft MCP GatewayMCP onlyEntra ID auth with role-based access to registered serversKubernetes, session-aware routingMIT
IBM ContextForgeMCP, A2A and REST federationUser-scoped access, OAuth and JWT; virtual servers wrapping REST APIsPython, Docker, KubernetesApache 2.0
agentgatewayHTTP, gRPC, MCP and A2A data planeTool access policies per MCP target; JWT and OIDCLinux Foundation project, Kubernetes-orientedOpen source
LiteLLMLLM and MCP proxyMCP server access by key, team or organisationPython proxyMIT core, enterprise add-ons

Bifrost is the top pick for teams running agents in production, for three reasons that follow from its documentation. It defaults to deny at the key level and enforces the list at both tools/list and tools/call, which is the posture OWASP’s excessive-agency guidance asks for. Its per-user OAuth, per-user headers and token-exchange modes address the credential problem more thoroughly than a single shared upstream token. And because the same virtual key governs model routing, budgets and tool access, a team gets one identity model rather than two gateways to reconcile. Code Mode is the differentiator for teams with hundreds of tools; the Bifrost MCP gateway overview collects the relevant documentation in one place.

The dedicated proxies have real strengths. Docker’s container isolation directly answers the local-server compromise risk the specification describes, Microsoft’s gateway suits teams standardised on Entra ID and Kubernetes, and ContextForge fits when internal REST services need to become MCP tools without new code. For how the LLM side of these products compares, see the five LLM gateways compared on this site, the survey of leading AI gateways for production, and the review of open-source LLM gateways you can self-host, which also notes which ones treat MCP as a native feature and which as a plugin.

When to add an MCP gateway, and how to start

The trigger is sharing, not size. A single developer with one or two local servers gains little from a gateway. The case becomes strong when any of the following is true:

  • More than one team or product uses the same MCP servers.
  • Any MCP server holds write access to production systems, customer data or source code.
  • Agents run unattended, in CI or as background workers, with no human watching each tool call.
  • The combined tool list is large enough that definitions take a visible share of every prompt; in the Code Mode benchmark above, classic tool calls already carried a large overhead at 96 tools.
  • Security or compliance needs to answer “which agent called which tool, as whom” after the fact.

A sensible rollout order, keyed to risk rather than features:

  1. Inventory first. List the MCP servers in use, including those on developer machines. The gap between the servers a platform team knows about and those actually configured is usually the first finding.
  2. Put write-capable servers behind the gateway. Start with the servers that can change things: source control, ticketing, databases, cloud consoles.
  3. Issue per-agent keys with explicit allow-lists. Start from deny and add tools as each agent needs them. Resist * except for read-only servers.
  4. Move credentials to the gateway. Prefer per-user OAuth or token exchange where the upstream supports it; retire tokens stored in local configs.
  5. Turn on logging and review it. Tool-call logs catch misbehaving agents; policy-change audit logs catch misconfigured keys.
  6. Add guardrails and code execution where the data justifies them. Argument guardrails for destructive tools; Code Mode once tool counts climb.

For teams comparing self-hosted AI gateway options more broadly, the deciding question is whether MCP governance should live in the same gateway as model routing. The better choice for most agent platforms is yes, because the identity that pays for a model call should be the same identity that is allowed to call a tool.

An MCP gateway will not make an unsafe tool safe. What it does is make the question “what can this agent reach?” answerable in one place, and make the answer smaller. Teams evaluating an MCP gateway can request a Bifrost demo or start from the open-source repository.

Feature claims are drawn from each vendor’s public documentation as of September 2026.

Sources

  1. MCP specification 2025-11-25: Transports Model Context Protocol
  2. MCP specification 2025-11-25: Authorization Model Context Protocol
  3. MCP specification 2025-11-25: Tools Model Context Protocol
  4. MCP security best practices Model Context Protocol
  5. OWASP MCP Top 10 (2025, beta) OWASP Foundation
  6. MCP03:2025 Tool Poisoning OWASP Foundation
  7. LLM06:2025 Excessive Agency OWASP GenAI Security Project
  8. MCP Security Notification: Tool Poisoning Attacks Invariant Labs
  9. GitHub MCP Exploited: Accessing private repositories via MCP Invariant Labs
  10. Code execution with MCP Anthropic
  11. Bifrost MCP overview and gateway mode Maxim AI
  12. Bifrost MCP tool filtering per virtual key Maxim AI
  13. Bifrost tool filtering levels Maxim AI
  14. Bifrost Virtual MCPs Maxim AI
  15. Bifrost MCP authentication Maxim AI
  16. Bifrost Code Mode Maxim AI
  17. Bifrost MCP Code Mode benchmark report Maxim AI
  18. Bifrost guardrails (including MCP guardrails) Maxim AI
  19. Bifrost observability (LLM and MCP logs) Maxim AI
  20. Bifrost audit logs Maxim AI
  21. Bifrost Edge: govern MCP servers Maxim AI
  22. Docker MCP Gateway Docker
  23. Microsoft MCP Gateway Microsoft
  24. IBM ContextForge IBM
  25. agentgateway Linux Foundation
  26. LiteLLM MCP gateway BerriAI

Frequently asked questions

What is an MCP gateway?

An MCP gateway is a proxy that sits between AI agents and the Model Context Protocol servers they use. Agents connect to one endpoint; the gateway authenticates the caller, returns only the tools that caller is allowed to use, injects the credential for the upstream server, applies checks to each tool call and logs it. It turns many point-to-point tool connections into one governed path.

What is the difference between an MCP server and an MCP gateway?

An MCP server exposes tools, such as reading a GitHub issue or querying a database, and executes them. An MCP gateway exposes no tools of its own; it aggregates many servers behind one endpoint and decides who can reach which tool, under which credential, with what logging. The server is the execution layer and the gateway is the policy layer in front of it.

Is there an open-source MCP gateway?

Yes. Bifrost (Apache 2.0) combines an MCP gateway with an LLM gateway. Docker MCP Gateway runs servers as containers, Microsoft MCP Gateway targets Kubernetes, IBM ContextForge federates MCP, A2A and REST, agentgateway is a Linux Foundation data plane, and LiteLLM includes an MCP gateway in its proxy. They differ mainly in policy depth and deployment model.

Do I need an MCP gateway if I only use one or two MCP servers?

Usually not. A single developer with one or two local servers gains little from a gateway beyond logging. The case strengthens once several agents or people share the same servers, once credentials for production systems are involved, or once the combined tool list is large enough that its definitions take a noticeable share of every prompt.

How does an MCP gateway reduce token usage?

In two ways. Filtering removes tools a caller does not need from the list the model sees, so fewer definitions are sent on each turn. Code execution patterns, such as Bifrost's Code Mode, replace hundreds of tool definitions with a few meta-tools the model uses to load signatures on demand and run a short script, keeping intermediate results out of the context window.

Can an MCP gateway prevent tool poisoning?

Partly. A gateway can keep unreviewed servers out of the tool list, pin each caller to a curated set and run guardrails on tool arguments and results. It cannot tell a malicious instruction hidden in an approved tool's description from a legitimate one unless someone reviews those descriptions. Treat it as blast-radius reduction, not a detector.

How do you secure MCP servers?

Follow the MCP specification's authorisation rules for HTTP servers, validate token audience and never pass client tokens through. Grant each agent the minimum tools and scopes it needs, keep a human approval step for destructive tools, log every call, review tool descriptions before approval, and inventory local servers on developer machines so none run outside policy.

All tools →