ToolsComparison

Ten open-source AI gateways you can self-host in 2026, ranked

Bifrost, LiteLLM, agentgateway, Agent Router, APISIX, Higress, Kong, MLflow, Plano and New API ranked on licence, what the free tier holds back, footprint, MCP, governance and air-gapped fit.

Illustration of a server rack labelled your rack, your VPC with ten numbered gateway units, the first lit in a blue-violet gradient; a paper licence tag reading Apache 2.0, MIT, OSI hangs beside it, and a card shows a crossed-out cloud with the words air-gapped, no phone-home.
Illustration of a server rack labelled your rack, your VPC with ten numbered gateway units, the first lit in a blue-violet gradient; a paper licence tag reading Apache 2.0, MIT, OSI hangs beside it, and a card shows a crossed-out cloud with the words air-gapped, no phone-home.

An AI gateway is a reverse proxy that understands model APIs: applications send requests to one endpoint with one credential, and the gateway picks the provider, applies budgets and rate limits, retries failures and records what each call cost. For a lot of organisations the gateway also has to run on their own hardware, because prompts carry customer data, contracts specify where that data may go, or the network has no route to the internet at all. That narrows the field to open-source gateways you can operate yourself, and it changes which questions matter.

This ranking covers ten gateways that ship under an OSI-approved licence and run on your own infrastructure, assessed as of October 2026. It is written for platform and AI engineers choosing a gateway for a VPC, an on-premises cluster or an air-gapped network. Every product claim comes from the project’s own documentation, repository or announcements, read in the first week of October 2026, and every performance figure is labelled with the conditions under which its publisher measured it.

The headline: Bifrost, an open-source AI gateway written in Go by Maxim AI, ranks first because it puts governance, an MCP gateway and observability in the free build and publishes the lowest overhead figures. LiteLLM is second on breadth and familiarity. The Kubernetes-native projects agentgateway and Agent Router follow. The API-gateway plugin approach (APISIX, Higress, Kong) suits teams that already run those gateways, with the caveat that Kong now reserves most of its AI features for a paid licence. Readers who also want hosted options should start with the companion ranking of AI gateways in 2026, and readers new to the category with what an AI gateway is.

How we ranked open-source AI gateways

The list admits a gateway only if two things are true. Its gateway code is published under an OSI-approved licence (Apache 2.0, MIT or AGPL-3.0 here), and a team can run it entirely on its own infrastructure without a vendor-hosted control plane. A product whose AI features require a commercial licence to function at all, such as Gravitee’s LLM proxy, whose documentation lists an Enterprise licence as a prerequisite, is excluded however good its paid edition is.

Within that pool, each gateway was assessed on ten criteria:

  1. Licence and the open-core line. What the open-source build includes, and what the vendor holds back for an enterprise edition. This weighs most, because a self-hosted gateway whose governance is paid is a trial, not a choice.
  2. Footprint and language. What you run: one binary, a Python service with PostgreSQL and Redis, or a Kubernetes control plane with an Envoy data plane.
  3. Performance overhead. Published latency and throughput, with conditions. Few projects publish anything; Bifrost’s published benchmarks are used as the reference point.
  4. Provider coverage. How many model providers are native, and whether OpenAI-compatible endpoints fill the gaps.
  5. Governance. Virtual keys, budgets, token-based rate limits, RBAC and SSO, and which tier each sits in.
  6. Guardrails. PII, secrets and content checks on the request path, and whether they need a cloud service.
  7. Observability. OpenTelemetry, Prometheus, request logs and cost attribution.
  8. MCP support. Whether the gateway can sit in front of MCP servers with authentication and tool-level access control.
  9. Kubernetes and air-gapped deployment. Helm charts, Gateway API support, and whether anything phones home.
  10. Community activity. Release cadence, maintenance and foundation governance, checked on 5 October 2026.

Three projects that appear on many older lists did not make it. TensorZero’s repository was archived on 12 June 2026 and the project describes itself as no longer maintained, although its Apache 2.0 code remains available. kgateway, the CNCF Sandbox gateway, moved its AI and MCP features into agentgateway from version 2.3.0 and now focuses on being an Envoy API gateway. Gravitee is excluded on licence, as above.

The contenders

RankGatewayLicenceLanguage and footprintDeploymentBest fit
1BifrostApache 2.0; enterprise features paidGo, single binary; SQLite or PostgreSQLBinary, Docker, Helm, in-VPCTeams that want governance, MCP and low overhead in one service
2LiteLLMMIT; enterprise/ separately licensedPython; PostgreSQL, Redis beyond one instanceDocker, Helm, TerraformPython-first teams that need the widest provider list
3agentgatewayApache 2.0; Solo Enterprise editionRust data planeBinary, Docker, Kubernetes Gateway APIPlatform teams governing agents, MCP and A2A traffic
4Agent Router (formerly Envoy AI Gateway)Apache 2.0Go control plane, Envoy data planeKubernetes with Envoy Gateway; aigw locallyKubernetes teams already on Envoy Gateway
5Apache APISIXApache 2.0; API7 Enterprise adds managementLua on OpenResty; etcd or standalone YAMLDocker, Helm, ingress controllerTeams already running APISIX
6HigressApache 2.0; Alibaba Cloud managed editionGo control plane, Envoy and IstioDocker all-in-one, HelmTeams moving off ingress-nginx, Wasm plugin authors
7Kong Gateway OSSApache 2.0; most AI plugins need an AI LicenseLua on OpenResty; DB-less or PostgreSQLDocker, Kubernetes ingressExisting Kong estates with modest AI needs
8MLflow AI GatewayApache 2.0Python, inside the MLflow servermlflow serverTeams already running MLflow tracking
9Plano (formerly Arch)Apache 2.0 proxy; routing models under Katanemo licenceRust and Envoy with Wasm pluginsNative, Docker, KubernetesAgent apps that want model-based routing and filter chains
10New APIAGPL-3.0 with attribution termsGo; SQLite, MySQL or PostgreSQL; optional RedisDocker, multi-nodeInternal key and quota distribution with a built-in console

Three columns: purpose-built AI gateways (Bifrost, LiteLLM, MLflow, New API) with keys, budgets and logs built in; Envoy and Rust data planes (agentgateway, Agent Router, Plano) as Kubernetes Gateway API resources routing MCP and A2A; and API-gateway plugins (Kong, APISIX, Higress) that add AI to existing routes

Figure 1: The ten gateways fall into three camps, and the camp matters more than any single feature because it decides what you operate.

The ten open-source AI gateways, ranked

1. Bifrost

Bifrost is an AI gateway written in Go and published under Apache 2.0, with a release (HTTP v2.2.5) on 2 October 2026. It starts with npx -y @maximhq/bifrost or a single Docker container, serves an OpenAI-compatible API and a web UI on the same port, and stores configuration in SQLite or PostgreSQL. The supported-providers matrix lists more than 30 providers, from OpenAI, Anthropic, Bedrock, Vertex AI and Azure to Groq, Mistral, Ollama, vLLM and SGLang, plus custom provider instances. The LLM gateway overview summarises the product for buyers.

Strengths. The open-source build carries the parts most vendors charge for. Virtual keys set deny-by-default provider and model access, token and request rate limits, and budgets across a customer, team and key hierarchy, with resets from one minute to one year; a request is rejected if any tier is over its limit. Fallbacks, weighted load balancing, semantic caching (Redis or Valkey, Weaviate, Qdrant or Pinecone, default similarity threshold 0.8), Prometheus metrics, OpenTelemetry tracing and Go or WASM plugins are all listed as open-source in the documentation overview. The MCP gateway acts as both MCP client and server over STDIO, HTTP and SSE, filters tools per virtual key, supports per-user OAuth, and offers a Code Mode that the documentation says cuts input tokens by up to 92.8 per cent.

On performance, Bifrost’s published benchmarks report 11 µs of gateway overhead at 5,000 requests per second on a t3.xlarge (59 µs on a t3.medium) with a 100 per cent success rate. Its head-to-head with LiteLLM on a t3.medium at 500 requests per second reports a P99 of 1.68 s against 90.72 s and 424 against 44.84 requests per second. These are vendor figures with a public harness.

Limitations. Guardrails (Bedrock Guardrails, Azure Content Safety, Model Armor, Presidio, secrets detection and others), clustering, adaptive load balancing, OpenID Connect SSO, RBAC, audit logs, the Datadog connector and log exports come with the enterprise tier, so teams that need multi-node high availability or those controls should plan for the enterprise licence.

Best for: teams that want one self-hosted service to handle routing, budgets, MCP and telemetry at high request rates, and enterprises that need in-VPC or air-gapped deployment with a supported upgrade path to SSO, audit and guardrails.

2. LiteLLM

LiteLLM is one of the most widely adopted open-source gateways, with v1.104.0 released on 3 October 2026. It is MIT-licensed except for an enterprise/ directory under its own licence, and it is a Python SDK as well as a proxy server. The README claims “100+ LLM providers”, the widest native coverage on this list.

Strengths. Virtual keys, per-key and per-team budgets, spend tracking, fallbacks and an OpenTelemetry callback are free. The MCP gateway supports Streamable HTTP, SSE and stdio with access control by key, team or organisation, and nothing in its documentation is marked enterprise-only. Guardrail integrations including Presidio, Lakera, Bedrock, Azure Content Safety and OpenAI Moderation are in the free tier. Docker images, a Helm chart and Terraform modules for AWS and GCP exist, and configuration is the YAML many Python teams already know.

Limitations. The enterprise feature list holds back SSO beyond five users, SCIM, JWT and OIDC auth, audit logs, secret-manager integrations, tag-based budgets, and per-key or per-model guardrails. Virtual keys and stored MCP servers need PostgreSQL, and the production guidance calls for Redis once there is more than one instance. Security has been a recurring theme in 2026: on 24 March, compromised PyPI packages 1.82.7 and 1.82.8 were live for about 40 minutes, which the project’s incident write-up attributes to a compromised CI/CD dependency (the official Docker image was unaffected); CVE-2026-42208, a SQL injection in API-key checking, was fixed in 1.83.7; and a host-header authentication bypass was fixed in 1.84.0. Pin versions and track advisories. Teams weighing a move can compare options in this site’s LiteLLM alternatives guide.

Best for: Python-first teams that need the broadest provider list and are prepared to run PostgreSQL and Redis alongside the proxy.

3. agentgateway

agentgateway is a Rust data plane, contributed by Solo.io to the Linux Foundation in August 2025 and accepted as the fourth hosted project of the Agentic AI Foundation in June 2026. It released v1.6.0 on 2 October 2026. The AAIF announcement cites more than 300 contributors from over 60 organisations, including Microsoft, Red Hat and Salesforce.

Strengths. It treats LLM, MCP and agent-to-agent (A2A) traffic as one routing problem. MCP federation works over all transports, A2A is supported, and the gateway is Gateway API conformant on Kubernetes while also running as a standalone binary or container. It fronts more than 20 providers through an OpenAI-compatible API, supports JWT, OIDC, API keys, mTLS and virtual keys, attributes cost in US dollars per request, and applies CEL-based authorisation. Guardrails cover regex, OpenAI moderation and AWS and Google services. Rate limits can count tokens locally or call a remote Envoy rate-limit service, with maxTokens per period to enforce spend limits per key. Telemetry is OpenTelemetry metrics, logs and traces. The project’s own v1.3 benchmark (Fortio, 32 connections, 1 KB payloads, mock backend) reports a P99 of 1.97 ms at about 36,900 queries per second.

Limitations. Solo Enterprise for agentgateway adds an advanced UI, on-behalf-of identity, FIPS and long-term-support builds, and advertises global request- and token-based rate limiting; check which build covers the limit model you need. There is no hierarchical budget model comparable to Bifrost’s or LiteLLM’s, and air-gapped operation is not documented, although nothing in the standalone binary requires egress.

Best for: platform teams that want one gateway for models, MCP servers and agent-to-agent calls, configured as Kubernetes resources.

4. Agent Router (formerly Envoy AI Gateway)

The project known as Envoy AI Gateway was renamed Agent Router on 10 September 2026 and moved from the CNCF’s Envoy umbrella to the Agentic AI Foundation. The announcement is explicit that the code, maintainers, release cadence, Apache 2.0 licence, aigateway.envoyproxy.io CRDs and aigw CLI are unchanged. Its v1.0.0 reached general availability on 23 June 2026 and v1.1.0 shipped on 21 August.

Strengths. It builds on Envoy Gateway, so AI routes are Kubernetes resources and the data plane is Envoy. Provider credentials stay in the gateway, failover is built in, and it supports 16 providers according to the release notes. Usage-based rate limiting charges tokens after the response completes and lets a CEL expression weight input, cached and output tokens differently, keyed on user and model headers. The MCP gateway, with tool routing and authorisation, has been in the open-source build since v1.0, and v1.1 added OpenTelemetry GenAI semantic conventions and a Grafana dashboard. aigw run starts an OpenAI-compatible router locally on port 1975 for development.

Limitations. Production use assumes Kubernetes 1.32 or later, Envoy Gateway 1.8.1 or later and Redis for global limits, which is a lot of machinery if you do not already run it. The pages read for this article document no native content guardrail and no budget hierarchy, and no performance figures are published.

Best for: Kubernetes platform teams already standardised on Envoy Gateway who want AI routing and MCP as cluster resources.

5. Apache APISIX

Apache APISIX is a Lua gateway on OpenResty, with release 3.19.0 on 28 September 2026. Its AI support is a set of plugins: ai-proxy and ai-proxy-multi for routing, ai-rate-limiting, ai-cache, ai-rag, prompt guard, decorator and template plugins, content moderation through AWS, Aliyun or Lakera, and openapi-to-mcp.

Strengths. Nothing AI-related appears to be held back. API7’s edition comparison says both editions include the AI plugins, with the enterprise product adding console RBAC, SSO, audit logging, multi-cluster management and Vault secrets. ai-proxy-multi supports OpenAI, Azure OpenAI, Anthropic, Gemini, Vertex AI, Bedrock, DeepSeek, OpenRouter and any OpenAI-compatible endpoint, with weighted, consistent-hash and semantic routing, health checks and fallback on 429 and 5xx. ai-rate-limiting limits by total, prompt or completion tokens per consumer or per model instance, with Redis, Redis Cluster or Sentinel for shared counters. Prometheus exports LLM latency including time to first token and token counts. Helm charts and an ingress controller cover Kubernetes, and standalone YAML mode removes etcd.

Limitations. Governance is the gateway’s general consumer model, not an AI-specific hierarchy of teams and budgets. The older mcp-bridge plugin is deprecated. And API7 announced AISIX in March 2026, a separate Rust AI gateway, because building AI on plugins was “pushing the limits of the underlying architecture”; that is a signal about where the vendor’s AI investment is going.

Best for: teams that already run APISIX and want AI routes beside their existing APIs with nothing held back by licence.

6. Higress

Higress is an Envoy- and Istio-based gateway that started at Alibaba, now lives at higress-group/higress, and joined the CNCF Sandbox in March 2026. Release v2.2.5 on 4 October 2026 moved to Envoy 1.36.10 with 29 CVE fixes.

Strengths. The AI feature set is broad: token-based rate limiting, multi-model fallback, model-aware routing, semantic caching, RAG and protocol conversion that the project says covers more than 100 models. It is a conformant Gateway API Inference Extension implementation. Plugins are Wasm and can be written in Go, Rust or JavaScript. MCP servers can be hosted through plugins, and an openapi-to-mcp converter turns existing APIs into tools. It starts with a single Docker command or a Helm chart, and the CNCF post positions it as a path off ingress-nginx.

Limitations. The built-in ai-security-guard plugin calls Alibaba Cloud Content Security and needs an access key, so the shipped guardrail does not work in an air-gapped network; you would write or import your own. Much of the community, documentation and plugin ecosystem is Chinese-first. A2A exists only as a feature request, and there is no AI-specific budget hierarchy.

Best for: teams replacing ingress-nginx, or comfortable writing Wasm plugins, who want AI traffic management in the same gateway.

7. Kong Gateway OSS with AI plugins

Kong Gateway is Apache 2.0, written in Lua on OpenResty, with one of the largest API-gateway communities. Its open-source repository contains six AI plugins: ai-proxy, ai-prompt-guard, ai-prompt-decorator, ai-prompt-template, ai-request-transformer and ai-response-transformer.

Strengths. For an existing Kong estate, AI routes reuse key-auth, JWT, OAuth2, ACLs, request-count rate limiting and the Prometheus, OpenTelemetry and logging plugins that teams already operate. The 3.9 drivers cover OpenAI, Azure, Anthropic, Cohere, Mistral, Llama, Hugging Face, Gemini and Bedrock, and the gateway runs DB-less or with PostgreSQL, in hybrid control-plane and data-plane mode, or behind the Kubernetes ingress controller.

Limitations. The open-source line has stalled. The last open-source release is 3.9.3, from 17 June 2026, while the documentation shows features badged for versions up to 3.16 that exist only in Kong’s commercial builds. ai-proxy-advanced, ai-semantic-cache, ai-rate-limiting-advanced (token-based limits), ai-sanitizer, ai-aws-guardrails and the AI MCP Proxy all carry an “AI License Required” badge. The free guardrail is regex allow and deny lists. Self-hosting Kong for AI on the open-source licence means a basic proxy without MCP, semantic caching or token limits.

Best for: organisations already running Kong OSS that need simple provider proxying, or that intend to buy the AI License.

8. MLflow AI Gateway

MLflow, the Linux Foundation machine-learning platform, with release 3.16.1 in September 2026, rebuilt its gateway in early 2026; the introduction post dates native provider integrations to 3.11 and lists more than 15 providers. The gateway runs inside mlflow server and is Apache 2.0.

Strengths. Every request becomes an MLflow trace with the full payload, latency and token counts, beside the experiments and evaluations teams already keep there. The documentation covers encrypted central credentials, traffic splitting for A/B tests, fallback chains, usage tracking and guardrails that use an LLM judge to block or redact requests, with preset safety and PII instructions. Budget policies set a US dollar threshold per day, week or month that alerts or rejects, optionally per workspace, with a Redis tracker for multiple replicas.

Limitations. Authentication is HTTP basic auth, and budget administration needs authentication enabled. The gateway is a feature of a Python ML platform rather than a dedicated proxy, no overhead figures are published, MCP is not described, and the LLM-judge guardrails add a model call to each checked request.

Best for: data-science and ML platform teams already running MLflow who want gateway traces to land next to their evaluation data.

9. Plano (formerly Arch)

Plano, Katanemo’s “AI-native proxy server and data plane for agentic apps”, is the project formerly called Arch gateway; the old archgw repository now redirects to it. It is built on Envoy by Envoy contributors, mostly in Rust, Apache 2.0, with release 0.4.37 on 28 September 2026.

Strengths. Routing can be by model name, by semantic alias or automatically by preference, using Katanemo’s small routing models (Plano-Orchestrator-4B by default, Arch-Router-1.5B). Filter chains add jailbreak protection, moderation and memory as in-process MCP filters or HTTP filters. It falls back on 429 and 5xx, emits OpenTelemetry traces and metrics, and runs natively via planoai up, in Docker or on Kubernetes according to the deployment guide.

Limitations. The routing models are not OSI-licensed: Arch-Router is distributed under the Katanemo licence, and the hosted copies are free only in a US-central region, so an air-gapped deployment must run them locally. No virtual keys, budgets or per-tenant limits are documented, no Helm chart is documented, and MCP is used for filters rather than as a tool gateway.

Best for: teams building agent applications that want preference-based model routing and guardrail filter chains at the proxy, with governance handled elsewhere.

10. New API

New API is a Go gateway with a React console, built on the MIT-licensed One API and itself licensed AGPL-3.0, with release v1.0.0-rc.41 on 30 September 2026. Its community is concentrated in China.

Strengths. It is built for distributing model access inside an organisation: channels with priorities, weights and retries; per-key restrictions; quotas and subscriptions; expression-based pricing; OAuth and OIDC login with two-factor authentication; and rate limits per node or shared through Redis. It speaks OpenAI Chat and Responses, Anthropic Messages and Gemini formats with streaming and tools. A single Docker container with SQLite is enough to start, with MySQL or PostgreSQL, an optional ClickHouse log store and multi-node deployment for scale.

Limitations. The AGPL-3.0 licence carries an added requirement that modified versions keep a visible attribution and link in the user interface, which legal teams should review. The README does not document MCP, OpenTelemetry, Prometheus, guardrails or a Helm chart, and it warns operators running public or resale services to meet local licensing and content-safety rules.

Best for: teams that want an internal portal for handing out model keys with quotas and a billing view, more than a policy engine.

What actually differs between self-hosted AI gateways

Feature lists converge; four things do not.

Where the free tier stops

The licence on the repository is only half the answer. Bifrost’s open-source build includes keys, budgets, MCP, caching and telemetry, and charges for guardrails, clustering, SSO and audit. LiteLLM’s free tier is similar, with SSO free for five users. agentgateway’s open-source build is generous, with the enterprise edition adding operational extras. Kong is the outlier: the open-source build is frozen at 3.9.3 and the AI features most teams want carry an AI License requirement. APISIX holds nothing AI-related back.

Four rows comparing what the open-source licence includes and what the paid edition adds: Bifrost (keys, budgets, MCP, OTel, cache free; guardrails, clustering, RBAC, audit paid), LiteLLM (keys, budgets, MCP, OTel free; SSO past five users, audit, per-key guardrails paid), Kong Gateway (OSS 3.9.3 with ai-proxy and regex prompt guard; AI License for MCP, semantic cache, token limits), agentgateway (MCP, A2A, guardrails, OTel free; Solo Enterprise for UI, FIPS, LTS, on-behalf-of identity)

Figure 2: For a self-hosted gateway, the open-core line decides whether the free build is a product or a trial.

Governance depth

Only Bifrost and LiteLLM ship an AI-specific hierarchy of customers or organisations, teams and keys with budgets that roll up, which is what lets a platform team give fifty internal teams their own keys and spending caps. agentgateway, Agent Router, APISIX and Higress offer token-based rate limits keyed on consumers or headers, which controls load but not spend across a hierarchy. MLflow has workspace budgets. New API has quotas per key. For the wider control picture of identity, budgets and audit, Maxim AI’s AI governance overview describes how Bifrost layers SSO, SCIM and audit logs onto the open-source virtual keys. Bifrost applies those governance and security controls (virtual keys, budgets, guardrails, audit logs) centrally at the gateway, and Bifrost Edge extends the same policies to AI traffic from desktop apps, browser tools and coding agents on employee machines, with endpoint enforcement on each device. That matters for the shadow AI problem a server-side gateway cannot see.

MCP and agent traffic

MCP has become a gateway requirement in 2026. Bifrost, LiteLLM, agentgateway, Agent Router and Higress route MCP in their free builds; agentgateway also handles A2A. The details differ: per-key tool filtering, per-user OAuth and whether the gateway can aggregate many servers behind one endpoint. This site’s explainer on MCP gateways covers the mechanics, and the sibling ranking of open-source MCP gateways compares them in depth.

Observability and air-gapped fit

Most of the ten speak OpenTelemetry; Bifrost, APISIX and Higress also export Prometheus natively. Maxim AI’s AI observability page sets out the gateway-level view (logs, metrics, traces and cost per provider, model and team, collected off the hot path) that Bifrost implements. Air-gapped fit depends on hidden dependencies more than on the gateway: Higress’s shipped guardrail calls Alibaba Cloud, Plano’s default routing models are hosted, and LLM-judge guardrails in MLflow need a reachable model. Bifrost documents in-VPC deployments with no external network dependencies.

Recommendations by constraint

Decision tree from self-hosting a gateway: already running Kong, APISIX or Higress leads to adding the AI plugins after checking the free tier; Kubernetes-first with agents and MCP leads to agentgateway or Agent Router; needing budgets and low overhead leads to Bifrost, or LiteLLM for Python-first teams; already on MLflow leads to MLflow AI Gateway

Figure 3: The estate you already run settles most self-hosting decisions before features do.

  • No gateway estate yet, many teams, real traffic. Start with Bifrost. It is one binary, it carries budgets and MCP in the free build, and its published overhead is the lowest on the list. Plan for the enterprise licence if you need clustering, SSO, guardrails or audit logs.
  • Python-first team, widest provider list. LiteLLM, with PostgreSQL and Redis provisioned from the start, pinned versions and a habit of reading its security advisories.
  • Kubernetes platform team, agents and MCP. agentgateway if you want MCP, A2A and LLM traffic in one Rust data plane; Agent Router if you already run Envoy Gateway.
  • You already run APISIX, Higress or Kong. Add AI routes to what you operate. APISIX holds nothing back; Higress needs a local guardrail for air-gapped use; Kong’s free tier is a basic proxy unless you buy the AI License.
  • ML platform on MLflow. MLflow AI Gateway keeps traces next to evaluations, at the cost of a thinner proxy.
  • Strict air gap. Choose a gateway whose guardrails and routing models run locally, and test with egress blocked before go-live.

For a narrower shortlist read from the same angle, Maxim AI’s own guide to the five best open-source LLM gateways for self-hosted deployments and its production-ready comparison of the top LLM gateways cover overlapping ground, and GenAI Brief’s five-gateway comparison goes deeper on failover semantics.

The bottom line

Self-hosting a gateway is a commitment to operate it for years, so the licence and the open-core line matter more than this month’s feature list. The safest choice is a gateway whose free build already covers the controls you need, whose community is active and whose dependencies you can run offline. On those terms, Bifrost leads in October 2026, LiteLLM remains the broad default for Python shops, and the Kubernetes-native projects are the ones to watch as agent traffic grows. Before committing, run your own load profile through the shortlist, block egress, and check which features your licence actually includes.

Feature claims are drawn from public documentation and vendor announcements as of October 2026; check current docs before deciding.

Sources

  1. Bifrost repository and Apache 2.0 licence Maxim AI
  2. Bifrost benchmarks Maxim AI
  3. Bifrost documentation overview (open-source and enterprise features) Maxim AI
  4. Bifrost virtual keys Maxim AI
  5. Bifrost MCP gateway overview Maxim AI
  6. Bifrost guardrails Maxim AI
  7. Bifrost clustering Maxim AI
  8. Bifrost Edge overview Maxim AI
  9. LiteLLM enterprise features BerriAI
  10. LiteLLM MCP gateway BerriAI
  11. LiteLLM security update, March 2026 BerriAI
  12. Apache APISIX ai-proxy-multi plugin Apache Software Foundation
  13. Apache APISIX ai-rate-limiting plugin Apache Software Foundation
  14. Kong AI MCP Proxy plugin Kong
  15. Kong Gateway repository Kong
All tools →