ModelsComparison
Top 10 open-weight models in October 2026, ranked for teams that self-host
MiMo-V2.6-Pro, GLM-5.3, DeepSeek V4.1 and seven more, ranked on independent evals, licence terms, active parameters, hardware and serving support, each read from its own model card.

An open-weight model is one whose trained parameters you can download, run on your own hardware and fine-tune, whatever the terms that come with them. In October 2026 the best of these models sit within a few months of the closed frontier, and the choice among them is no longer about capability alone. It is about which licence your lawyers will sign, how many GPUs the weights need, and whether vLLM and SGLang run the model well on the day you need it.
This ranking is for teams that self-host or fine-tune: platform engineers sizing a cluster, ML engineers choosing a base for post-training, and engineering leads deciding which workloads leave the hosted APIs. It covers the latest downloadable releases from Xiaomi, Z.ai, DeepSeek, Moonshot, Alibaba’s Qwen team, NVIDIA, Google, MiniMax, Mistral and OpenAI, read from each model card and licence file in the first week of October 2026. Independent standing comes from the Artificial Analysis Intelligence Index (version 4.3.2) and the LMArena text leaderboard (updated 2 October 2026). For the closed side of the market, see our companion ranking of the top AI models in October 2026; for why the open models trail by months rather than years, see open weights trail the frontier by months.
The short version: MiMo-V2.6-Pro is the strongest open model and is MIT-licensed. GLM-5.3 and DeepSeek V4.1 Flash are the practical workhorses. Kimi K3 is the most capable agent on vendor benchmarks but the hardest to host and the most conditional to license. Below 100 billion parameters, Qwen3.8-27B is in a class of its own.
How we ranked
Six criteria, applied in roughly this order of weight:
- Independent evaluation standing. The Artificial Analysis Intelligence Index aggregates ten evaluations, including agentic, long-context and knowledge tasks, run by a third party. LMArena’s text leaderboard is a human-preference signal from more than 8.6 million votes. Vendor benchmark tables are reported, attributed and treated as claims. Our guide to reading a model card explains why.
- Licence terms. MIT, Apache 2.0 and OpenMDW-1.1 count as truly open for commercial use. Licences that add revenue thresholds, attribution clauses or model-as-a-service gates lose ground in proportion to how many readers they will affect.
- Parameters and hardware. Total parameters set the memory bill; active parameters set the per-token compute. A model that fits one eight-GPU node beats an equally capable one that needs two.
- Context window. Native length, and whether long context is a trained property or an extension.
- Tool use. Native function calling, agentic benchmark results and controllable reasoning effort.
- Ecosystem support. Whether vLLM, SGLang and llama.cpp (or Ollama and LM Studio) are documented on the card, and whether quantised variants exist.
Hardware notes in the table are our arithmetic from the published parameter counts and precisions, for weights alone. The KV cache, activations and batching headroom come on top, and they scale with context length and concurrency.
The ten at a glance
| Rank | Model | Params (total / active) | Licence | Context | Hardware note (weights) | Best fit | Source |
|---|---|---|---|---|---|---|---|
| 1 | MiMo-V2.6-Pro | 1.02T / 42B | MIT | 1M | ~1 TB in FP8; one 8-GPU Blackwell node | Highest open capability, no strings | Card |
| 2 | GLM-5.3 (and Flash) | 753B / 40B (Flash 320B / 18B) | GLM-5.3 licence (Flash MIT) | 1M | ~750 GB in FP8; one 8 x H200 node | Coding and long-horizon agents | Card |
| 3 | DeepSeek V4.1 Flash | 552B / 16B decode | MIT | 1M | ~550 GB in FP8; one 8-GPU node | Throughput per GPU, long inputs | Card |
| 4 | Kimi K3 | 2.8T / 104B | Kimi K3 licence | 1M | ~1.4 TB in MXFP4; multi-node on Hopper | Agentic and browsing workloads | Card |
| 5 | Qwen3.8 (27B and 2.4T-A95B) | 27B dense; 2.4T / 95B | Apache 2.0 (27B); Qwen3.8-Max licence | 262K native, 1M extended | 55.6 GB BF16 (27B); ~4.8 TB BF16 (2.4T) | Best single-GPU model; fine-tuning base | Card |
| 6 | Nemotron 3 Ultra | 550B / 55B | OpenMDW-1.1 | 1M | 8 x B200 or GB300 per card; NVFP4 | Full openness, NVIDIA stacks | Card |
| 7 | Gemma 4 | 2.3B to 30.7B; 26B / 3.8B MoE | Apache 2.0 | 256K | Phone to one GPU | Edge, on-device, multimodal | Card |
| 8 | MiniMax-M3 | 428B / 23B | MiniMax Community License | 1M | ~430 GB in FP8; one 8-GPU node | Long-context multimodal coding | Card |
| 9 | Mistral 3 family | 675B / 41B (Large 3); 128B dense (Medium 3.5); 119B / 6.5B (Small 4) | Apache 2.0 (Large, Small); modified MIT (Medium) | 256K | 8 x H200 in FP8 (Large 3) | European vendor, Apache base | Card |
| 10 | gpt-oss-120b | 117B / 5.1B | Apache 2.0 | 131K | One 80 GB GPU in MXFP4 | Cheapest capable reasoning model | Card |
Independent scores, for reference: on the Artificial Analysis index v4.3.2, MiMo-V2.6-Pro scores 46, GLM-5.3 45, Kimi K3 44, GLM-5.3 Flash 42, Qwen3.8-2.4T-A95B 40, DeepSeek V4.1 Flash 39, Qwen3.8-27B 34, MiniMax-M3 29, Nemotron 3 Ultra 23, Gemma 4 31B 15, Mistral Medium 3.5 14, gpt-oss-120b 12 and Mistral Small 4 11. On LMArena’s text leaderboard, Kimi K3 is the highest open model at 1488 (16th overall), then MiMo-V2.6-Pro at 1480, GLM-5.3 at 1478 and DeepSeek V4.1 Flash at 1474. The overall leader, a proprietary Google model, sits at 1525.

Figure 1: Hardware sorts the field before benchmarks do. Most teams will choose within one rung, not across all four.
1. MiMo-V2.6-Pro (Xiaomi)
MiMo-V2.6-Pro is Xiaomi’s flagship reasoning model, released on 21 September 2026 and published on Hugging Face as MiMo-V2.6-Pro-RL. It is a sparse mixture-of-experts model with 1.02 trillion total and 42 billion active parameters, 384 routed experts with eight active per token, and 70 layers split between 60 sliding-window and ten global-attention layers. It ships with a 681-million-parameter vision encoder and a multi-token-prediction module for speculative decoding, and supports a 1M-token context.
Strengths. It is the best open model by independent measurement: first among open-weight models on the Artificial Analysis index at 46, and the top MIT-licensed model on LMArena at 1480. The licence is plain MIT, with no revenue thresholds and no attribution clause, which makes it the strongest model in this list that a reseller or fine-tuner can use without reading further. The card reports 89.9 on Terminal-Bench 2.1, 82.0 on OSWorld-Verified, 76.9 on Toolathlon-Verified and 71.9 on DeepSWE v1.1, all vendor-run.
Limitations. It is large and slow. Artificial Analysis measured 41.2 output tokens per second and flagged high verbosity, at 140 million output tokens to complete its evaluation suite, which inflates the serving bill per task. The card’s SGLang example uses 16-way tensor parallelism with two data-parallel replicas, and the vLLM example uses eight-way tensor parallelism; in FP8 the weights alone are about 1 TB, so on Hopper hardware this is a two-node model and on Blackwell a full node. The --trust-remote-code flag in both examples means custom modelling code, which some security reviews will want to read. A smaller MiMo-V2.6-Flash-RL exists for teams that want the family without the footprint.
Best fit: teams that want the most capable open model available, under the least restrictive licence, and can dedicate a Blackwell node to it.
2. GLM-5.3 and GLM-5.3 Flash (Z.ai)
GLM-5.3 is Z.ai’s August 2026 release, dated 18 August by Artificial Analysis. It reuses the GLM-5.2 base and improves through extended post-training, with 753 billion total and 40 billion active parameters, a mixture-of-experts design with dynamic sparse attention, and a 1M-token context. Reasoning effort is selectable per request (low, high, max). GLM-5.3 Flash, the smaller sibling, has 320 billion total and 18 billion active parameters, is the first natively multimodal model in the GLM-5 series, and is MIT-licensed.
Strengths. The pair takes second and fourth place among open models on the Artificial Analysis index (45 and 42) and sits just behind MiMo on LMArena (1478 and 1473). Z.ai positions GLM-5.3 as its strongest coding model and reports 66.9 on DeepSWE, 78.1 on FrontierSWE and 84.5 on CyberGym; Flash reports 84.3 on Terminal-Bench 2.1. Ecosystem support is the broadest at this size: the card documents vLLM, SGLang, TokenSpeed, KTransformers, Unsloth and Transformers, Ascend NPU paths through vLLM-Ascend, xLLM and SGLang, and quantised builds for llama.cpp, Ollama and LM Studio. In FP8 the full model fits a single 8 x H200 node; Flash fits comfortably on 8 x H100.
Limitations. GLM-5.3 left MIT. Its licence requires any licensee running a model-as-a-service business whose group revenue exceeds $10 billion over 12 months to pass Z.ai’s security review before use. That affects very few readers, but it is a condition, and LMArena’s leaderboard still labels the model MIT, which is a reminder to read the licence file rather than a summary. The CyberGym and ExploitBench results also mean the model is more capable at offensive security work than most, which some risk teams will want to assess.
Best fit: coding agents and long-horizon engineering tasks on one node, with Flash as the cost-efficient, MIT-licensed default.
3. DeepSeek V4.1 Flash (DeepSeek)
DeepSeek V4.1 Flash, released on 10 September 2026, is an unusual design: a causal encoder-decoder with a 20-layer encoder and a 20-layer decoder, each layer a mixture of experts with one shared and 384 routed experts. It has 552 billion backbone parameters and activates 8 billion per token during prefill and 16 billion during decode. It accepts images and text, generates text, supports a 1M-token context and is MIT-licensed. The card says it was trained from scratch on 45 trillion multimodal tokens.
Strengths. It is the efficiency leader among frontier-class open models. The card reports a global KV cache of 890 bytes per token, about a quarter of DeepSeek V4 Flash’s, using a second-generation compressed sparse attention and FP4 caching. That matters more than parameter count for long-context serving, because KV cache is what limits concurrency at 1M tokens. Artificial Analysis measured 213 output tokens per second, the fastest of the leading open models, and scored it 39; LMArena puts it at 1474. Vendor-reported agentic results are strong: 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1. vLLM, SGLang, Transformers and Docker Model Runner are documented.
Limitations. There is no V4.1 Pro yet; the larger V4 Pro line (1.6 trillion total, 49 billion active, MIT) is a generation older. The encoder-decoder architecture is new, so serving-engine optimisations and quantised community builds lag the decoder-only models, and fine-tuning recipes are less settled. Nine points behind MiMo on the Artificial Analysis index is a real gap on the hardest tasks.
Best fit: high-volume, input-heavy workloads (retrieval, document processing, code review across large repositories) where cost per token and concurrency matter more than the last few points of capability.
4. Kimi K3 (Moonshot AI)
Kimi K3, released on 16 July 2026, is the largest model in this list: 2.8 trillion total parameters with 104 billion active, routing each token through 16 of 896 experts. It uses Moonshot’s Kimi Delta Attention and attention residuals, was trained with quantisation awareness to MXFP4 weights and MXFP8 activations, and supports a 1,048,576-token context. The card recommends vLLM, SGLang and TokenSpeed.
Strengths. It is the most preferred open model by human raters: 1488 on LMArena, 16th overall and above every other open release. On the Artificial Analysis index it scores 44, third among open models. Its vendor-reported agentic numbers are the strongest published by any open lab: 94.5 on MCPMark-Verified, 91.2 on BrowseComp, 84.8 on OSWorld-Verified and 88.3 on Terminal-Bench 2.1, with 93.5 on GPQA Diamond. Native 4-bit training means the published weights are the intended precision rather than a lossy post-hoc quantisation.
Limitations. Hardware and licence. Even at MXFP4, 2.8 trillion parameters is roughly 1.4 TB of weights by our arithmetic, which means multi-node serving on Hopper and a full high-memory Blackwell node at minimum. With 104 billion active parameters per token it is also expensive to run: Artificial Analysis measured 45 output tokens per second. The Kimi K3 licence is MIT-style with two additions: any product above 100 million monthly active users or $20 million in monthly revenue must display “Kimi K3” prominently, and any model-as-a-service business above $20 million in revenue over 12 months needs a separate agreement with Moonshot. That second clause is aimed at inference hosts and will catch more readers than GLM-5.3’s $10 billion threshold.
Best fit: well-resourced teams building browsing, computer-use and MCP-heavy agents for internal use, where the licence’s service clause does not bite.
5. Qwen3.8-27B and Qwen3.8-2.4T-A95B (Alibaba)
The Qwen3.8 generation arrived in August 2026 as two very different open releases. Qwen3.8-2.4T-A95B, published on 12 August, is the first Max-class Qwen with open weights: 2.4 trillion total and 95 billion active parameters, 512 experts with 11 active, a hybrid of Gated DeltaNet linear attention and gated full attention, a native 262,144-token context extendable to 1,010,000, and thinking mode that cannot be disabled. Qwen3.8-27B, published on 14 August, is a dense 27-billion-parameter vision-language model under Apache 2.0 with the same native and extended context.
Strengths. Qwen3.8-27B is the strongest small open model by a wide margin. It scores 34 on the Artificial Analysis index, more than double Gemma 4 31B’s 15 and nearly triple gpt-oss-120b’s 12, and the card reports 61.7 on SWE-bench Pro, 73.0 on Terminal-Bench 2.1 and 84.3 on OSWorld-Verified. Its 55.6 GB of BF16 weights fit one 80 GB GPU unquantised, and dense weights are the easiest to fine-tune with standard tooling. The 2.4T flagship scores 40 and reports 67.7 on SWE-bench Pro.
Limitations. The flagship is hard to justify for most self-hosters. Its card publishes BF16 weights only, which is about 4.8 TB by our arithmetic, Artificial Analysis measured it at 37.7 tokens per second, and it scores below the far smaller GLM-5.3 and MiMo. Its Qwen3.8-Max licence requires a separate commercial licence for organisations above $50 million in annual revenue running model-as-a-service or AI work-assistant businesses (coding or office tools), plus prominent attribution above 100 million monthly users or $20 million in monthly revenue. LMArena lists the hosted Qwen3.8 Max as proprietary, a further sign that the two should not be confused.
Best fit: Qwen3.8-27B is the default single-GPU model and the best base for domain fine-tuning in 2026.
6. Nemotron 3 Ultra (NVIDIA)
Nemotron 3 Ultra, released on 4 June 2026, has 550 billion total and 55 billion active parameters in a latent mixture-of-experts design that interleaves Mamba-2 state-space layers, MoE layers and a few attention layers, with multi-token prediction. It was pre-trained in NVFP4 and supports a 1M-token context according to the card. It is the only model here released under the Linux Foundation’s OpenMDW-1.1, which covers weights, code, documentation and data in one instrument.
Strengths. Openness. NVIDIA published large portions of the 20-trillion-token pre-training corpus and its post-training collections alongside the weights, which no other model in this list does at this scale. For regulated buyers who need training-data provenance, and for researchers who want to reproduce or extend the recipe, that is decisive. At launch, Artificial Analysis called it the leading US open-weight model and measured it at over 400 output tokens per second, faster than its usual serving of gpt-oss-120b despite being more than four times larger. The card documents vLLM, SGLang and TensorRT-LLM, with NVFP4 weights for Blackwell.
Limitations. Capability has been overtaken. On the current v4.3.2 index it scores 23, half of MiMo-V2.6-Pro, and the card’s 56.4 on Terminal-Bench 2.1 trails the Chinese models’ high 80s. The hardware guidance is 8 x B200 or GB300 for a single node, or multi-node on H100 and H200; TensorRT-LLM support is Blackwell-only. It is the right model for reasons other than leaderboard position.
Best fit: organisations that need documented training data under a permissive licence, or that are standardised on NVIDIA’s NIM and TensorRT-LLM stack.
7. Gemma 4 (Google)
Gemma 4, announced on 2 April 2026, is Google’s open family and the first Gemma under the OSI-approved Apache 2.0 licence rather than Google’s custom terms. It spans E2B (2.3 billion effective parameters), E4B, a 12B model, a 26B mixture of experts with 3.8 billion active, and a 31B dense model at 30.7 billion. The larger models accept text and images with a 256K context; audio is supported on E2B, E4B and 12B.
Strengths. Range and ecosystem. Gemma 4 runs from phones to a single server GPU, has native function calling, and the 31B card documents Transformers, vLLM, SGLang and Docker Model Runner, with Ollama, llama.cpp and LM Studio through quantisations. Google reports more than 400 million community downloads and over 100,000 variants across the Gemma line. The 31B card reports 85.2 on MMLU Pro, 89.2 on AIME 2026 and 80.0 on LiveCodeBench v6, and LMArena places it at 1453. At launch, Artificial Analysis noted Gemma 4 31B scored slightly above Nemotron 3 Ultra on its coding sub-index.
Limitations. At the top size it now trails Qwen3.8-27B badly on the independent index (15 against 34), and Artificial Analysis notes it is slower than comparable models. It is a generation behind Gemini by design.
Best fit: on-device and edge deployments, multimodal assistants on consumer hardware, and teams that want a Google-maintained Apache 2.0 base.
8. MiniMax-M3 (MiniMax)
MiniMax-M3, released on 1 June 2026, has 428 billion total and 23 billion active parameters, a 1M-token context and native multimodal training on images and video. Its MiniMax Sparse Attention is the headline: the card claims 9x faster prefill and 15x faster decode than M2 at 1M context, with per-token compute cut to a twentieth. Thinking can be enabled, adaptive or disabled per request.
Strengths. Long-context efficiency and coding. The card reports 80.5 on SWE-bench Verified and 59.0 on SWE-bench Pro, 78.1 on MMMU Pro and 85.4 on Video-MME v2. It scores 29 on the Artificial Analysis index, above Nemotron 3 Ultra and every Western open model, at 23 billion active parameters. Serving support covers SGLang, vLLM, Transformers, KTransformers and Unsloth, with ATOM for MXFP4 and MXFP8 quantisation. The card publishes BF16 weights and points to ATOM for MXFP8 and MXFP4; at 8-bit the weights fit one 8 x H100 node.
Limitations. The licence is the most restrictive in this list. The MiniMax Community License allows free non-commercial use; any commercial use, explicitly including deploying a fine-tuned or otherwise modified model, requires displaying “Built with MiniMax M3”, a one-time notice to MiniMax below $20 million in annual revenue, and prior written authorisation above it. It also carries a prohibited-use appendix that includes military applications. On LMArena it sits lower than its index score suggests, at 1440.
Best fit: research, internal tools and smaller companies that need efficient 1M-token multimodal context and accept the attribution and notice terms.
9. Mistral 3 family (Mistral AI)
Mistral’s open line in October 2026 is three models. Mistral Large 3 (December 2025) has 675 billion total and 41 billion active parameters, a 2.5-billion-parameter vision encoder, a 256K context and an Apache 2.0 licence. Mistral Small 4 (March 2026) has 119 billion total and 6.5 billion active, also Apache 2.0. Mistral Medium 3.5 is a dense 128-billion-parameter multimodal model with a 256K context under a modified MIT licence.
Strengths. A European vendor with Apache 2.0 weights and serious commercial backing: Mistral raised a €3 billion round in September 2026 and in August opened regional EU and US inference endpoints. Large 3’s card gives concrete deployment targets, FP8 on a single 8 x H200 node or NVFP4 on a single H100 or A100 node, and native function calling. Medium 3.5’s card documents vLLM, SGLang, Transformers, Ollama, LM Studio and llama.cpp, configurable reasoning effort, and vendor-reported results of 77.6 on SWE-Bench Verified and 91.4 on τ³-Telecom.
Limitations. Mistral’s open models trail on independent measurement: Medium 3.5 scores 14 and Small 4 scores 11 on the Artificial Analysis index, and on LMArena Large 3 sits at 1414 and Medium 3.5 at 1427. Large 3’s card notes Transformers support was not ready at release. Medium 3.5’s modified MIT licence requires companies above $20 million in monthly revenue to buy a commercial licence or use Mistral’s hosted service. Mistral itself now hosts third-party open models, starting with GLM-5.2, which says something about where the capability is.
Best fit: European organisations that want a domestic vendor relationship, Apache 2.0 weights and the option of the same models on Mistral’s regional API.
10. gpt-oss-120b (OpenAI)
gpt-oss-120b, released on 5 August 2025, is OpenAI’s open reasoning model: 117 billion total and 5.1 billion active parameters, a 131K context, three reasoning-effort levels and an Apache 2.0 licence. Its sibling gpt-oss-20b fits in 16 GB.
Strengths. It is the cheapest capable reasoning model to run. MXFP4 quantisation of the MoE weights lets it fit one 80 GB H100 or MI300X, and Artificial Analysis measured 196 output tokens per second. Fourteen months on, its ecosystem is the most mature here: the card lists vLLM, SGLang, Transformers, Ollama, LM Studio and llama.cpp, and it supports function calling, web browsing and structured outputs. Full chain-of-thought access helps debugging.
Limitations. It is the oldest model in the list and shows it: 12 on the Artificial Analysis index, far behind Qwen3.8-27B at 34. It requires OpenAI’s harmony response format to work correctly, which trips up generic chat templates, and its knowledge cutoff is mid-2024.
Best fit: cost-sensitive reasoning on a single GPU, classification and extraction at volume, and teams already tooled for it.
Close calls and what we left out
Meta’s Muse Glimmer, released on 10 August 2026 under Apache 2.0, is a 30-billion-parameter multimodal agent model that Meta says runs on 24 or 32 GB consumer GPUs and is supported in Ollama, LM Studio, llama.cpp, vLLM and SGLang. It scores 17 on the Artificial Analysis index, above Gemma 4 31B and gpt-oss-120b, and a reader who wants a Meta model on a laptop could reasonably swap it in at tenth. We kept gpt-oss for its single-GPU throughput and the depth of its tooling. Meta’s Llama 4 models remain available under the Llama 4 Community License but have been overtaken; Meta’s frontier work now happens in the closed Muse Spark, with Glimmer distilled from it.
MiMo-V2.6-Flash, DeepSeek V4 Pro and IFM’s K2 Horizon (31 on the index) were also considered. Each is credible, but each is beaten within its own size and licence tier by a model above.
What actually differs
Licences sort into three tiers
The permissive tier (MIT, Apache 2.0, OpenMDW-1.1) covers MiMo-V2.6-Pro, DeepSeek V4.1 Flash, GLM-5.3 Flash, Qwen3.8-27B, Nemotron 3 Ultra, Gemma 4, Mistral Large 3 and Small 4, and gpt-oss. The threshold tier adds conditions that bite only above a revenue line or for inference resellers: GLM-5.3 at $10 billion, Kimi K3 at $20 million of service revenue, Qwen3.8-Max at $50 million, Mistral Medium 3.5 at $20 million a month. MiniMax-M3 is alone in requiring notice for any commercial use. The pattern from our earlier analysis holds: labs that need the weights to pay tighten the terms on their best model while keeping the smaller one permissive.

Figure 2: The licence decides who can use a model before the benchmark decides who should.
Active parameters set the bill, total parameters set the cluster
Six of the ten activate between 16 and 55 billion parameters per token, which is why they generate quickly once loaded. The cost is memory: every expert must be resident. The practical boundary is one eight-GPU node. GLM-5.3, DeepSeek V4.1 Flash, MiniMax-M3 and Mistral Large 3 fit one Hopper node in FP8; MiMo-V2.6-Pro and Nemotron 3 Ultra want Blackwell; Kimi K3 and Qwen3.8-2.4T do not fit on one Hopper node at any published precision.
Long context is now the norm, KV cache is the constraint
Seven of the ten advertise 1M tokens. Few teams will serve that length at concurrency, because KV cache grows with every token in flight. DeepSeek’s 890 bytes per token and MiniMax’s sparse attention are the two published answers; for the others, plan context budgets from the cache, not the card.
Serving support has converged on vLLM and SGLang
Every model here documents vLLM and SGLang, and TokenSpeed now appears on the Qwen, Kimi and GLM cards. llama.cpp, Ollama and LM Studio support is the real differentiator for workstation use: documented for Gemma 4, gpt-oss, Mistral Medium 3.5, Muse Glimmer and quantised GLM-5.3, and absent from the cards of the trillion-parameter models.
Serving open weights alongside hosted models
Few teams will run only open weights. The common pattern is a self-hosted model for steady, high-volume or data-sensitive traffic, with a hosted frontier model for the hardest prompts and as a fallback when the GPU pool is saturated. That puts a routing decision in front of every request, and it is the job an LLM gateway is built for.
Bifrost, an open-source AI gateway written in Go by Maxim AI, is one way to do this from a single self-hosted binary. It treats self-hosted inference servers as providers in their own right: vLLM, SGLang, Ollama and Hugging Face are addressed as vllm/<model>, sgl/<model>, ollama/<model> and huggingface/<model>, alongside hosted providers such as OpenAI, Anthropic, Bedrock, Gemini and Mistral, and any OpenAI-compatible server can be added as a custom provider. For vLLM, the provider configuration sets the server URL and loaded model per key, so two GPU pools serving different models appear as two keys.
Three capabilities matter for a mixed open and hosted estate:
- Fallbacks to hosted models. A request can carry an ordered fallback list, for example a self-hosted GLM-5.3 first and a hosted model second. Each provider gets its full retry budget; fallbacks trigger only on retryable errors, not on validation failures, and the response records which provider served it.
- Weighted routing and model allowlists. Virtual keys carry
provider_configswith a weight per provider andallowed_modelswith regex support, so a team can be sent 80 per cent to self-hosted DeepSeek V4.1 Flash and 20 per cent to a hosted model, or restricted to MIT and Apache models only. - Governance and budgets. Budgets sit at customer, team and virtual-key level with reset windows from a minute to a year, plus token and request rate limits per key. The AI governance controls apply identically whether a request lands on your GPUs or on a provider, which matters when self-hosted inference has a real cost per GPU-hour that finance wants attributed.
Overhead is small enough not to matter next to model latency: Bifrost’s published benchmarks report 11 µs of gateway overhead at 5,000 requests per second on a t3.xlarge. Beyond the data centre, Bifrost Edge extends the same virtual keys, budgets, guardrails and audit logs to AI traffic on employee machines, with endpoint enforcement on each device.
Bifrost is not the only option. LiteLLM and other gateways can front vLLM too; a comparison of open-source LLM gateways for self-hosted deployments covers the trade-offs, as does our own comparison of five LLM gateways. Some features, including clustering and in-VPC deployment, are part of Bifrost’s enterprise tier.

Figure 3: A gateway turns “self-host or use an API” from an architecture decision into a routing rule you can change per team.
Recommendations by constraint
You need the most capable open model and no licence conditions. MiMo-V2.6-Pro, on a Blackwell node. If that is too large, GLM-5.3 Flash under MIT is the next step down.
You have one eight-GPU Hopper node. DeepSeek V4.1 Flash for throughput and long inputs, GLM-5.3 for coding agents. Both fit in FP8 with room for KV cache.
You have one GPU. Qwen3.8-27B. It is the strongest model under 100 billion parameters by a wide margin and Apache 2.0. Choose gpt-oss-120b if raw tokens per second matter more than capability.
You resell inference. Stay in the permissive tier: MiMo, DeepSeek, GLM-5.3 Flash, Qwen3.8-27B, Nemotron, Gemma, Mistral Large 3 or gpt-oss. Kimi K3 above $20 million in service revenue and MiniMax-M3 at any commercial scale need agreements first.
You need training-data provenance. Nemotron 3 Ultra is the only frontier-scale option with released pre- and post-training data under one licence.
You are building on-device or edge features. Gemma 4 E2B to 12B, or Muse Glimmer on a 24 GB card.
You want European vendor continuity. Mistral Large 3 under Apache 2.0, with the same family on Mistral’s regional API as a hosted fallback.
You want open and hosted models behind one API. Put a self-hosted gateway in front of both, set the open model as primary and the hosted model as fallback, and attribute spend by team from day one.
The bottom line
The open-weight field in October 2026 is deep enough that the best choice depends more on your hardware and your licence than on the leaderboard. MiMo-V2.6-Pro is the strongest model you can download and use without conditions. DeepSeek V4.1 Flash and GLM-5.3 are what most platform teams should benchmark first on their own workloads. Qwen3.8-27B has made the single-GPU tier competitive in a way it was not a year ago. Read the licence file, size the KV cache, run your own evaluation on a few hundred real prompts, and keep a hosted model behind the same endpoint for the tail.
Feature claims are drawn from public documentation and vendor announcements as of October 2026; check current docs before deciding.
Sources
- MiMo-V2.6-Pro-RL model card Xiaomi, via Hugging Face
- GLM-5.3 model card Z.ai, via Hugging Face
- GLM-5.3 licence Z.ai, via Hugging Face
- GLM-5.3-Flash model card Z.ai, via Hugging Face
- Kimi K3 model card Moonshot AI, via Hugging Face
- Kimi K3 licence Moonshot AI, via Hugging Face
- DeepSeek-V4.1-Flash model card DeepSeek, via Hugging Face
- Qwen3.8-2.4T-A95B model card Qwen team, Alibaba, via Hugging Face
- Qwen3.8-Max licence Qwen team, Alibaba, via Hugging Face
- Qwen3.8-27B model card Qwen team, Alibaba, via Hugging Face
- NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 model card NVIDIA, via Hugging Face
- NVIDIA Nemotron 3 Ultra released Artificial Analysis
- Gemma 4 31B instruction-tuned model card Google, via Hugging Face
- Gemma 4: Expanding the Gemmaverse with Apache 2.0 Google Open Source Blog
- MiniMax-M3 model card MiniMax, via Hugging Face
- MiniMax Community License MiniMax, via Hugging Face
- Mistral Large 3 675B Instruct 2512 model card Mistral AI, via Hugging Face
- Mistral Medium 3.5 128B model card and licence Mistral AI, via Hugging Face
- gpt-oss-120b model card OpenAI, via Hugging Face
- Introducing Muse Glimmer: An Open Agentic Model Meta Superintelligence Labs
- Open source models compared on the Artificial Analysis Intelligence Index Artificial Analysis
- Text Arena leaderboard LMArena
- Bifrost supported providers Maxim AI
- Bifrost vLLM provider configuration Maxim AI
- Bifrost fallbacks Maxim AI
- Bifrost virtual keys Maxim AI
- Bifrost benchmarks Maxim AI
- Bifrost Edge overview Maxim AI
Frequently asked questions
What is the best open-weight model in October 2026?
On independent evaluations, Xiaomi's MiMo-V2.6-Pro. It ranks first among open-weight models on the Artificial Analysis Intelligence Index v4.3.2 with 46 and is the top MIT-licensed model on LMArena's text leaderboard. Kimi K3 scores higher on LMArena (1488 against 1480) but carries a more restrictive licence and needs roughly three times the active compute.
Which open-weight model can I run on a single GPU?
Qwen3.8-27B, Gemma 4 31B and gpt-oss-120b. Qwen3.8-27B is a dense 27B model with 55.6 GB of BF16 weights, so it fits one 80 GB GPU unquantised or a 32 GB card once quantised. gpt-oss-120b runs on one 80 GB GPU through MXFP4. Gemma 4 and Meta's Muse Glimmer target 24 to 32 GB consumer GPUs after quantisation.
Which open-weight models have truly permissive licences?
MIT covers MiMo-V2.6-Pro, DeepSeek V4.1 Flash and GLM-5.3 Flash. Apache 2.0 covers Gemma 4, Qwen3.8-27B, gpt-oss, Mistral Large 3, Mistral Small 4 and Muse Glimmer. NVIDIA's Nemotron 3 uses OpenMDW-1.1, which also covers its released training data. GLM-5.3, Kimi K3, Qwen3.8-2.4T, Mistral Medium 3.5 and MiniMax-M3 add conditions.
What is the difference between total and active parameters?
Mixture-of-experts models store many expert networks but route each token through only a few. Total parameters decide how much memory the weights need; active parameters decide how much compute each token costs. DeepSeek V4.1 Flash stores 552 billion parameters but uses 16 billion per decoded token, so it needs a full GPU node yet generates quickly.
Which open-weight model is best for agents and tool use?
Kimi K3 posts the strongest published agentic numbers, including 94.5 on MCPMark-Verified and 91.2 on BrowseComp. MiMo-V2.6-Pro and DeepSeek V4.1 Flash are close on terminal and software-engineering tasks, at 89.9 and 90.6 on Terminal-Bench 2.1. These are vendor-reported figures; the independent Artificial Analysis index puts MiMo first overall.
Can I use open-weight and hosted models behind the same API?
Yes. A self-hosted gateway exposes one OpenAI-compatible endpoint and routes each request to a vLLM or SGLang server or to a hosted provider. Bifrost supports vLLM, SGLang, Ollama and Hugging Face as providers next to hosted APIs, with per-request fallbacks and budgets on virtual keys.


