ToolsComparison

Top 10 AI gateways in 2026, ranked for production use

Bifrost, LiteLLM, Kong, Azure API Management, Agent Router, Vercel, Cloudflare, Databricks, OpenRouter and APISIX ranked on overhead, failover, governance, guardrails, MCP and deployment.

Illustration of a ranked leaderboard of ten gateway rows, the top row highlighted in a blue-violet gradient with a gold rank badge, beside a tilted chip reading 11 µs overhead and a tilted gradient card ticking off six ranking criteria from overhead to deployment.
Illustration of a ranked leaderboard of ten gateway rows, the top row highlighted in a blue-violet gradient with a gold rank badge, beside a tilted chip reading 11 µs overhead and a tilted gradient card ticking off six ranking criteria from overhead to deployment.

An AI gateway is a proxy that sits between applications and model providers and turns provider keys, retries, failover, budgets, guardrails and logging into configuration instead of code. By October 2026 most production AI teams run one, because the alternative is reimplementing the same rate-limit handling and spend tracking in every service. This ranking is for platform and AI engineers choosing a gateway for production traffic, managed or self-hosted, and it covers ten products that between them span every architectural option currently on the market.

The products have converged on a basic feature list: an OpenAI-compatible endpoint, some form of fallback, some form of key management. They still differ sharply on four things: how much latency they add under load, what they do with a 429 versus a 500, how deep their governance model goes, and where they can run. Those four differences decide most real selections, and they drive the ranking below.

The top pick is Bifrost, an open-source AI gateway written in Go by Maxim AI, which combines the lowest published overhead in the category with governance, guardrails and an MCP gateway in one binary. LiteLLM is second on breadth; Kong and Azure API Management follow for teams that already run an API-management platform. Every product claim below comes from the vendor’s own documentation or announcements, read between 1 and 5 October 2026, and every benchmark figure is attributed to whoever published it.

How this ranking was built

The ranking weights nine criteria, each read from primary documentation rather than marketing pages. A gateway scores well when the capability is documented precisely enough to configure without a sales call.

  1. Performance overhead. Published latency added by the gateway, with the conditions it was measured under. Hosted gateways are judged on architecture, since none publishes an overhead figure.
  2. Provider coverage. How many providers and modalities the gateway speaks, and whether existing provider SDKs work unchanged.
  3. Failover and routing. Whether the gateway distinguishes error classes (rate limit, bad credential, transient failure, client error) and what it does with each.
  4. Governance. Caller identity, budgets, rate limits, model allowlists, RBAC, SSO and audit logs.
  5. Guardrails. Content safety, PII redaction and prompt-injection checks, and which providers they integrate.
  6. Observability. Metrics, traces, request logs and cost attribution, and whether they use open formats such as OpenTelemetry.
  7. MCP support. Whether the gateway can govern Model Context Protocol tool traffic as well as model calls, and in which tier.
  8. Deployment options. Self-hosted, in-VPC, air-gapped, hosted, or tied to a cloud platform.
  9. Licence and tiering. What is open source, what is paid, and what has to run alongside the gateway.

Deployment and governance carry the most weight, because a gateway that cannot run where your data is allowed to go is disqualified regardless of features, and governance is the reason most organisations adopt a gateway in the first place. Readers new to the category may want to start with the site’s explainer on what an AI gateway is, and the five-way LLM gateway comparison goes deeper on failover semantics for the top contenders.

The ten contenders at a glance

RankGatewayType and licenceDeploymentBest fit
1BifrostPurpose-built gateway; Apache 2.0, paid enterprise tierSelf-hosted binary or container; in-VPC, clustering in enterpriseProduction teams wanting performance, governance and MCP in one gateway
2LiteLLMPurpose-built proxy; MIT with separately licensed enterprise directorySelf-hosted; PostgreSQL and Redis for multi-instancePython-first teams needing the widest provider list
3Kong AI GatewayPlugins on Kong Gateway; core Apache 2.0, most AI plugins enterpriseSelf-managed or KonnectOrganisations already running Kong
4Azure API Management AI gatewayManaged API-management service; proprietaryAzure-hosted, multi-region; containerised self-hosted gateway on Developer and Premium tiersMicrosoft and Azure-standardised enterprises
5Agent Router (formerly Envoy AI Gateway)Control plane on Envoy Gateway; Apache 2.0KubernetesPlatform teams already on Envoy and Kubernetes
6Vercel AI GatewayHosted gateway; proprietaryHosted onlyApplication teams wanting zero operations
7Cloudflare AI GatewayHosted at the edge; proprietaryHosted onlyTeams already on Cloudflare Workers
8Databricks Unity GatewayGovernance layer in Unity Catalog; proprietaryInside DatabricksData teams building agents on Databricks
9OpenRouterHosted model marketplace; proprietaryHosted onlyPrototyping and broad model access
10Apache APISIX AI gatewayPlugins on APISIX; Apache 2.0Self-hostedTeams already running APISIX

Grid of three columns: purpose-built gateways (Bifrost, LiteLLM), API-gateway extensions (Kong, Azure API Management, Agent Router, APISIX) and hosted or platform gateways (Vercel, Cloudflare, Databricks, OpenRouter), each column listing what that camp optimises for and what it gives up

Figure 1: The ten gateways fall into three camps, and the camp usually matters more than the individual product.

1. Bifrost

Bifrost is a purpose-built AI gateway written in Go and released under Apache 2.0, with a paid enterprise tier. It runs as a single binary or container, exposes an OpenAI-compatible endpoint, and accepts existing OpenAI, Anthropic, Bedrock and Google GenAI SDK calls by changing the base URL. Its supported-providers matrix lists about 30 providers, from OpenAI, Anthropic, Bedrock, Vertex AI and Azure to Groq, Mistral, Cerebras, Fireworks, Ollama and vLLM, and the LLM gateway overview claims access to more than 1,000 models through them. The project is active: release v2.2.5 of the HTTP transport shipped on 2 October 2026.

Strengths. Performance is the clearest. The published benchmarks report 11 µs of gateway overhead at 5,000 requests per second on an AWS t3.xlarge, and 59 µs on a t3.medium, both at a 100 per cent success rate. In the same page’s head-to-head against LiteLLM at 500 requests per second on a t3.medium, Bifrost held a P99 of 1.68 s against 90.72 s, throughput of 424 against 44.84 requests per second, and 120 MB of memory against 372 MB. Fallbacks are error-class-aware: a 429 rotates to the next key with backoff, a 401 rotates without waiting, a 5xx retries the same provider, and a 400 skips straight to the next fallback. Virtual keys carry budgets and rate limits in a customer, team and key hierarchy. The MCP gateway is both client and server, with tool filtering per virtual key and a Code Mode the docs say cuts input tokens by up to 92.8 per cent. Enterprise guardrails cover 16 providers with CEL rules, and gateway observability includes Prometheus metrics, OpenTelemetry traces and log exports.

Limitations. Guardrails, clustering, RBAC, SSO, audit logs and log exports come with the enterprise tier. Semantic caching connects to a vector store of the team’s choice.

Best for: the strongest all-round choice in this ranking, particularly for agent platforms and regulated teams that need governance depth, self-hosting and low overhead without stitching several tools together.

2. LiteLLM

LiteLLM is a Python proxy and SDK from BerriAI that calls “100+ LLMs” through an OpenAI-compatible interface. It is one of the most widely adopted open-source gateways, and its YAML configuration is familiar to most Python teams. The core is MIT-licensed, with an enterprise/ directory under a separate commercial licence.

Strengths. Provider breadth is unmatched among self-hosted options. The proxy offers virtual keys with per-key, per-team and per-user budgets, retry and fallback logic across deployments with cooldowns, typed fallback lists for context-window and content-policy errors, and several routing strategies. Its MCP gateway is in the open-source tier and controls tool access by key, team and organisation. Guardrail integrations are wide, observability callbacks reach most monitoring tools, and the project now covers A2A agents as well as model calls.

Limitations. 2026 has been a difficult security year. On 24 March, PyPI versions 1.82.7 and 1.82.8 were published with credential-stealing code for about 40 minutes, after the project’s CI pipeline was compromised through its Trivy dependency; the official proxy Docker image was not affected, and a clean v1.83.0 shipped on 30 March. The security advisory feed then lists a critical JWT authentication bypass (CVE-2026-35030), an SQL injection requiring a valid key (CVE-2026-42208), an MCP command injection requiring admin rights (CVE-2026-30623) and a host-header authentication bypass in June. All were patched, and the disclosures are thorough, but self-hosters need a disciplined upgrade process. Multi-instance deployments also need PostgreSQL and Redis alongside the proxy, and SSO beyond five users, audit logs and per-key guardrails are enterprise features.

Best for: Python-first teams that need the widest provider list and are prepared to run PostgreSQL, Redis and a tight patch cadence. Teams evaluating a move can compare options in the site’s LiteLLM alternatives guide.

3. Kong AI Gateway

Kong AI Gateway is a set of AI plugins on Kong Gateway, which itself is Apache 2.0. It handles LLM calls, MCP traffic and agent-to-agent traffic through one gateway, and it inherits Kong’s declarative configuration, consumer model and plugin ecosystem.

Strengths. The AI Proxy Advanced plugin offers seven load-balancing algorithms: round-robin, consistent hashing, least connections, lowest latency, lowest usage, semantic and priority. Failover criteria are explicit and opt-in per status code: client errors do not trigger failover unless codes such as http_429 are added. The AI MCP Proxy plugin converts REST APIs into MCP tools, proxies upstream MCP servers and aggregates tool sets behind consumer ACLs; as of AI Gateway 3.14 a denied MCP call returns HTTP 403 to match the MCP 2025-11-25 authorisation specification. Guardrail plugins integrate AWS Guardrails, Azure Content Safety and GCP Model Armor alongside regex and semantic prompt guards, and Konnect provides token and latency dashboards plus OpenTelemetry export.

Limitations. The features that make Kong a serious AI gateway are enterprise-only. Both AI Proxy Advanced and AI MCP Proxy state that they are “only available as part of our AI Gateway Enterprise offering”, and the MCP proxy needs Kong Gateway 3.12 or later. The free ai-proxy plugin targets one model per route. Kong publishes no AI-specific overhead benchmark, and semantic features need Redis or pgvector.

Best for: organisations already running Kong Enterprise or Konnect, for whom one policy model across REST, AI and MCP traffic outweighs a feature-by-feature comparison.

4. Azure API Management AI gateway

Microsoft’s AI gateway in Azure API Management is a set of capabilities inside the existing API Management service, not a separate product, and it applies to all service tiers with feature availability varying by tier. It governs OpenAI Chat Completions and Responses APIs, the Anthropic Messages API (v2 tiers), the Google Vertex AI API, models in Microsoft Foundry and providers such as Amazon Bedrock.

Strengths. Token governance is precise. The llm-token-limit policy sets tokens-per-minute limits or quotas over an hour, day, week, month or year on any counter key, and can pre-calculate prompt tokens to reject oversized requests before they reach the backend. llm-emit-token-metric sends token metrics with custom dimensions to Azure Monitor, and semantic caching uses llm-semantic-cache-store and llm-semantic-cache-lookup with Azure Managed Redis. The backend load balancer supports round-robin, weighted, priority and session-aware distribution, and its circuit breaker honours the backend’s Retry-After header, which suits teams balancing provisioned-throughput and pay-as-you-go deployments. API Management can expose REST APIs as MCP servers, govern existing MCP servers and import A2A agent APIs, and content safety runs through Azure AI Content Safety. Managed identities remove API keys for Azure backends, and on the Developer and Premium tiers the gateway component can run as a self-hosted container on Kubernetes or on premises, managed from Azure.

Limitations. The unified model API that exposes several providers behind one OpenAI-compatible endpoint is still in preview, as is the Foundry integration. Configuration is XML policy, which is powerful but verbose. Guardrails centre on Azure’s own content-safety service, and cost depends on the API Management tier rather than on traffic alone.

Best for: enterprises standardised on Azure and Microsoft Foundry that want AI traffic governed by the same API Management estate, identities and monitoring they already operate.

5. Agent Router (formerly Envoy AI Gateway)

Envoy AI Gateway changed name and home in September 2026. The Agentic AI Foundation announcement, dated 9 September, states that “Envoy AI Gateway is now Agent Router and is joining the Agentic AI Foundation.” The code, maintainers, CRDs, CLI and container images are unchanged, the licence remains Apache 2.0, and the project lists Bloomberg, Tetrate, Tencent Cloud and Nutanix among its public adopters.

Strengths. Agent Router is a control plane on Envoy Gateway, so it inherits Envoy’s data plane, which many platform teams already trust for ingress. Developers get an OpenAI-compatible API; platform teams manage providers, credentials, token quotas, routing and failover as Kubernetes resources. The 1.1 release added token-counting APIs across providers, per-request upstream credential overrides, stream idle timeouts with failover, hostname routing for the MCP gateway, and OpenTelemetry GenAI semantic conventions with a Grafana dashboard. A two-tier pattern separates a global entry gateway from fine-grained access to self-hosted models.

Limitations. It is Kubernetes-only, and running it well means running Envoy Gateway well. Budgets and governance are expressed as rate-limit and policy resources rather than a hierarchy of keys, teams and customers. The documentation does not describe built-in content guardrails, so those need an external filter. Documentation is thinner than for the purpose-built gateways.

Best for: platform teams already running Envoy and Kubernetes who want AI and MCP traffic governed through the same declarative stack.

6. Vercel AI Gateway

Vercel AI Gateway is a hosted gateway at ai-gateway.vercel.sh, available on all Vercel plans and callable from any infrastructure, not just Vercel deployments. It covers text, image, video, speech, transcription, realtime, embeddings and reranking, and accepts AI SDK, OpenAI Chat Completions, OpenAI Responses and Anthropic Messages requests.

Strengths. Vercel adds zero markup to provider token prices, including with bring-your-own-key. Provider ordering and model fallbacks are configurable, and request logs record every routing attempt with status, latency, tokens and cost, which makes it the best per-request failover trace among the hosted options. Budgets apply at team, project, API-key and team-member level, and the gateway supports coding agents through the Vercel CLI. Routing rules, regional inference and zero-data-retention controls cover common compliance asks.

Limitations. It cannot be self-hosted. Budgets are soft caps: the crossing request completes, and BYOK spend is metered separately and does not count towards those limits. If a BYOK request fails the gateway can fall back to Vercel’s system credentials, which teams with strict provider-account rules must disable deliberately. The documentation describes no MCP gateway and no content guardrails comparable to Bifrost’s or Kong’s.

Best for: application teams that want provider failover, budgets and good request traces without operating anything.

7. Cloudflare AI Gateway

Cloudflare AI Gateway runs on Cloudflare’s edge network, takes one line of code to adopt, and is available on all Cloudflare plans. It supports Workers AI, OpenAI, Anthropic, Google Gemini, Replicate and other providers.

Strengths. The feature list now spans caching, spend limits, rate limiting, an Auto Router, Dynamic Routing, custom costs, guardrails, data loss prevention, authentication, BYOK, analytics, logging and custom metadata. Dynamic Routing chains model nodes, conditional branches, percentage splits, rate-limit nodes and budget-limit nodes into a flow addressed as a model name, which handles A/B tests and gradual rollouts. For teams already on Workers, the gateway sits next to the rest of their stack, and it is available on every Cloudflare plan.

Limitations. It is hosted only. Caching is exact-match rather than semantic, and the September comparison found guardrails that do not support streamed responses and a single rate limit per gateway outside Dynamic Routing. Error handling does not distinguish a 429 from a 500; a failure advances the route. The documentation describes no MCP gateway.

Best for: teams on Cloudflare Workers that want caching, analytics and spend limits with minimal setup.

8. Databricks Unity Gateway

Databricks renamed its AI Gateway to Unity Gateway and made it generally available on 4 August 2026 as “the Databricks governance solution for enterprise AI, part of Unity Catalog.” It governs Databricks-hosted foundation models and, since 17 July 2026, external providers such as OpenAI, Anthropic and Amazon Bedrock through bring-your-own-key.

Strengths. Governance reuses Unity Catalog: models and MCP servers are registered as securable objects and granted with the same privileges as tables and volumes. The gateway enforces rate limits on model and MCP services, supports traffic splitting and fallbacks, and added Smart Routing on 13 August and spend limits on external providers on 28 August, according to the release notes. Inference tables log requests and responses to Delta tables, so usage analysis happens in the same lakehouse as everything else. MCP servers get tool filtering and service policies.

Limitations. It only makes sense inside Databricks. Applications outside the platform gain little, and the gateway cannot be deployed independently. Overhead is not published, and external providers only arrived in July 2026, so the BYOK path is newer than the rest of the product.

Best for: data and ML teams building agents on Databricks who want model and tool access governed by Unity Catalog.

9. OpenRouter

OpenRouter is a hosted marketplace that puts “hundreds of models from many providers” behind one OpenAI-compatible API. It is the fastest way to try a new model, and many teams use it as a provider behind another gateway; Bifrost, for example, lists OpenRouter as a supported provider.

Strengths. Default routing prioritises providers without significant outages in the previous 30 seconds, then picks among the cheapest, weighted by the inverse square of price. The provider object accepts order, allow_fallbacks, sort by price, throughput or latency, data_collection: "deny" and zdr: true to restrict requests to zero-data-retention endpoints. The FAQ states that prompts and completions are not logged by default.

Limitations. It is hosted only, and it is a reseller as much as a gateway: credit purchases by card carry a 5.5 per cent fee ($0.80 minimum), and BYOK usage carries a 5 per cent fee above a monthly free allowance. Governance does not reach the customer, team and key budget hierarchy of the self-hosted gateways, the documentation describes no content guardrails or MCP gateway, and enterprise residency needs depend on which upstream providers are selected.

Best for: prototyping, model exploration and small teams that value breadth over control.

10. Apache APISIX AI gateway

Apache APISIX adds AI plugins to an Apache 2.0 API gateway: ai-proxy, ai-proxy-multi, ai-rate-limiting, ai-prompt-guard, ai-cache, ai-rag and ai-lakera-guard, among others.

Strengths. The ai-proxy-multi plugin supports OpenAI, DeepSeek, Azure, Anthropic, OpenRouter, Gemini, Vertex AI, Amazon Bedrock and OpenAI-compatible APIs, balances load by weighted round-robin, consistent hashing or semantic similarity, falls back on 429s, 5xx errors and rate-limit exhaustion, and runs active health checks. Token-based rate limiting and response caching are built in, and everything is open source with no enterprise gate on the AI plugins.

Limitations. MCP support is weak: the mcp-bridge plugin, which bridges SSE clients to stdio MCP servers, is deprecated and not recommended for new deployments. There is no budget hierarchy or virtual-key model, guardrails are limited to prompt patterns and a Lakera integration, and the AI feature set is less documented than Kong’s equivalent.

Best for: teams already running APISIX who want AI traffic on the same gateway without a commercial licence.

What actually differs between the top ten

Feature tables flatten real differences. Five dimensions separate the gateways in practice.

Overhead and where the gateway runs

Only Bifrost publishes an overhead figure with its conditions, and it does so at two instance sizes plus a same-hardware head-to-head. The direction of the gap between a compiled Go binary and a Python proxy is unsurprising; the size is what matters for agent workloads, where one user action can trigger dozens of model calls. Hosted gateways add a network round trip that no vendor can state in advance, because it depends on where your application runs. API-gateway extensions inherit their host’s latency profile plus plugin work.

What happens on a 429

Bifrost, LiteLLM, Kong and APISIX distinguish rate limits from server errors, though in different ways: Bifrost rotates keys inside a provider before falling back, LiteLLM cools deployments down, and Kong and APISIX fail over on configurable status codes. Azure API Management’s circuit breaker honours Retry-After. Vercel, Cloudflare and OpenRouter advance an ordered list on any failure. All of these work; they need different configuration discipline, especially around monthly spend caps that return a 429 without a Retry-After header.

Governance depth

CapabilityBifrostLiteLLMKongAzure APIMAgent RouterVercelCloudflareDatabricksOpenRouterAPISIX
Per-caller keysVirtual keysVirtual keysConsumersSubscriptionsK8s policyAPI keysGateway tokenUnity CatalogAPI keysConsumers
Hierarchical budgetsCustomer, team, keyKey, team, userCost limits (ent.)Token quotasToken quotasTeam to memberBudget nodesPer-user capsLimitedNo
RBAC and SSOEnterpriseEnterpriseEnterpriseEntra IDKubernetes RBACVercel teamsCloudflare accountUnity CatalogLimitedLimited
Guardrails16 providers (ent.)Wide, partly ent.Cloud and custom (ent.)Content SafetyNot built inNot documentedBuilt inService policiesNot documentedPrompt guard, Lakera
MCP gatewayOpen sourceOpen sourceEnterpriseYesYesNoNoYesNoDeprecated

Bifrost and LiteLLM lead on budget hierarchy; Kong, Azure and Databricks lead on integration with an existing identity and policy estate. For a fuller treatment of how governance works at the gateway layer, and why it now has to reach endpoints too, see the site’s analysis of shadow AI governance. Bifrost is the one gateway in the ten that extends its controls beyond the server: Bifrost Edge runs on macOS, Windows and Linux and routes traffic from desktop AI apps, browser AI, coding agents and MCP servers through the gateway, so the same virtual keys, budgets, guardrails and audit logs apply on employee machines; the Edge overview lists rollout through Jamf, Intune and Kandji.

MCP as a first-class citizen

MCP support moved from novelty to requirement during 2026. Bifrost and LiteLLM put MCP gateways in their free tiers; Kong charges for it; Azure, Agent Router and Databricks treat MCP servers as governed resources inside their platforms. The hosted gateways have not followed. The companion ranking of MCP gateways in 2026 covers this dimension in depth, and the MCP gateway explainer covers the mechanism.

Licence and what you run alongside

Bifrost runs as one binary, plus a vector store if semantic caching is enabled. LiteLLM needs PostgreSQL and, for more than one instance, Redis. Kong and APISIX need their gateway data planes plus Redis or pgvector for semantic features. Agent Router needs Kubernetes and Envoy Gateway. The hosted options need nothing, but charge in different ways: Vercel adds no markup, OpenRouter charges on credit purchases and BYOK, and Azure charges by API Management tier. For a self-hosted shortlist only, the companion piece on open-source AI gateways narrows the field further.

Decision tree with four questions: residency or air-gap requirements lead to self-hosting, Bifrost first and LiteLLM for breadth; an existing API platform leads to extending Kong, Azure API Management or Agent Router; building on Databricks leads to Unity Gateway; wanting zero operations leads to Vercel or Cloudflare

Figure 2: One constraint usually settles the shortlist before any feature comparison begins.

Recommendations keyed to constraints

  • You have data-residency, in-VPC or air-gap requirements. Self-host Bifrost. It runs anywhere a container runs, keeps budgets, fallbacks and MCP in the open-source tier, and adds clustering, RBAC, SSO, audit logs and guardrails in the enterprise tier. LiteLLM is the alternative for Python-first teams willing to run PostgreSQL, Redis and a fast patch cadence.
  • You run agents at scale. Bifrost. Overhead per call, key rotation on 429, a retry budget per fallback provider and MCP tool filtering per virtual key are the things agent platforms exercise hardest. Validate the published numbers on your own instance sizes first, as the vendor’s own benchmarking guide advises.
  • You already run Kong or Azure API Management. Extend it. One policy model across all API traffic is worth more than marginal feature wins. Budget for Kong’s enterprise tier, and expect some Azure capabilities to remain in preview.
  • Your platform is Kubernetes and Envoy. Agent Router fits the stack you already operate, now with foundation governance behind it.
  • You build on Databricks. Unity Gateway keeps model and tool governance inside Unity Catalog, where your data permissions already live.
  • You want zero operations. Vercel AI Gateway for budgets, request traces and no markup; Cloudflare if you are on Workers. Plan the migration path to a self-hosted gateway before an enterprise customer asks for residency.
  • You are prototyping. OpenRouter for breadth, ideally behind an OpenAI-compatible interface so you can swap it out later.

Whichever you choose, keep the application talking to an OpenAI-compatible endpoint. Every gateway here speaks that dialect, which makes the gateway itself replaceable.

The verdict

The AI gateway category has split into three products that share a name: purpose-built control planes for model and tool traffic, AI extensions to existing API-management platforms, and hosted convenience layers. Bifrost leads the first camp and the ranking as a whole because it combines published, reproducible performance figures, explicit failure semantics, a governance hierarchy and an MCP gateway in an Apache 2.0 binary, and because it is the only option that extends the same policy to employee devices. LiteLLM remains a capable second for breadth, with a security record in 2026 that self-hosters should weigh. The API-management extensions win wherever the platform already exists, and the hosted gateways win wherever nobody wants to run anything.

Decide which of the three products you are buying before comparing features. For a second opinion, the Maxim AI team’s own production-ready comparison of LLM gateways and its review of open-source LLM gateways for self-hosted deployments cover overlapping ground from the vendor’s side. Then run the gateway you shortlist against your own traffic, because every benchmark in this article, including the best ones, was measured under someone else’s conditions.

Feature claims are drawn from public documentation and vendor announcements as of October 2026; check current docs before deciding.

Sources

  1. Bifrost benchmarks Maxim AI
  2. Bifrost benchmarking: getting started Maxim AI
  3. Bifrost repository and releases (Apache 2.0) Maxim AI
  4. Bifrost supported providers Maxim AI
  5. Bifrost fallbacks Maxim AI
  6. Bifrost virtual keys Maxim AI
  7. Bifrost MCP gateway overview Maxim AI
  8. Bifrost enterprise guardrails Maxim AI
  9. Bifrost Edge overview Maxim AI
  10. LiteLLM documentation BerriAI
  11. Security Update: Suspected Supply Chain Incident BerriAI
  12. LiteLLM security advisories BerriAI
  13. LiteLLM MCP gateway BerriAI
  14. Kong AI Proxy Advanced plugin Kong
  15. Kong AI MCP Proxy plugin Kong
  16. AI gateway capabilities in Azure API Management Microsoft
  17. Azure API Management self-hosted gateway overview Microsoft
  18. Vercel AI Gateway documentation Vercel
  19. Vercel AI Gateway budgets Vercel
  20. Cloudflare AI Gateway features Cloudflare
  21. Cloudflare AI Gateway Dynamic Routing Cloudflare
  22. Apache APISIX ai-proxy-multi plugin Apache Software Foundation
  23. Apache APISIX AI gateway Apache Software Foundation

Frequently asked questions

What is the best AI gateway in 2026?

For most production teams, Bifrost. It is open source under Apache 2.0, publishes 11 µs of overhead at 5,000 requests per second, classifies errors before retrying or falling back, and ships virtual keys, budgets and an MCP gateway in the free tier. Teams already standardised on Kong or Azure API Management should usually extend that platform instead.

What is the difference between an AI gateway and an API gateway?

An API gateway routes and secures HTTP traffic in general. An AI gateway understands model APIs: it counts tokens, prices requests, rotates provider keys, fails over between models and providers, and applies content guardrails. Kong, Azure API Management, Agent Router and APISIX add AI features to an API gateway; Bifrost and LiteLLM are built for model traffic from the start.

Which AI gateways can be self-hosted?

Bifrost, LiteLLM, Kong, Agent Router and Apache APISIX can all run on your own infrastructure. Azure API Management offers a containerised self-hosted gateway on its Developer and Premium tiers, still managed from Azure. Vercel AI Gateway, Cloudflare AI Gateway and OpenRouter are hosted services only, and Databricks Unity Gateway runs inside the Databricks platform.

Which AI gateways support MCP?

Bifrost and LiteLLM ship MCP gateways in their open-source tiers. Kong's AI MCP Proxy is enterprise-only and needs Kong Gateway 3.12 or later. Azure API Management can expose REST APIs as MCP servers and govern existing ones, Agent Router includes an MCP gateway, and Databricks governs MCP servers as Unity Catalog securables. Vercel, Cloudflare and OpenRouter do not document one.

Is OpenRouter an AI gateway?

It is a hosted model marketplace with gateway features: one OpenAI-compatible endpoint, automatic provider fallback, provider ordering and zero-data-retention filters. It does not offer self-hosting, team budgets at the depth of Bifrost or LiteLLM, or an MCP gateway, so it fits prototyping and model access better than enterprise governance.

How much latency does an AI gateway add?

It depends on the gateway and where it runs. Bifrost publishes 11 µs at 5,000 requests per second on an AWS t3.xlarge. Hosted gateways add a network round trip to their nearest point of presence, which depends on your region. Most vendors publish no comparable figure, so run your own load test before deciding.

All tools →