ModelsAnalysis

Open-weight models are closing the gap. The economics say it will not fully close

The lag between the best closed model and the best open-weight model has shrunk to months on most benchmarks. Whether it reaches zero depends on who pays for frontier training runs and why.

Illustration of a chart where a violet open-weight capability curve converges on a black closed-model curve, an unlocked padlock labelled weights public, and a tag reading gap months, not years.
Illustration of a chart where a violet open-weight capability curve converges on a black closed-model curve, an unlocked padlock labelled weights public, and a tag reading gap months, not years.

Three years ago the best open-weight model trailed the frontier by a year or more on nearly every benchmark. Today, on the standard suites, the lag is measured in months, and on price-performance the open models lead. That progress is real and it has changed what most teams should default to.

It has also produced a recurring prediction: that the gap will close entirely and closed models will become a premium brand for a commodity product. We think the prediction misreads where the gap comes from. It is not primarily technical. It is a question of who pays for frontier training runs and what they expect in return.

Where the gap is today

On knowledge, reasoning and coding benchmarks, leading open-weight releases from Chinese labs, from Meta, Mistral and a lengthening list of others have repeatedly matched closed models that were state of the art six to twelve months earlier. The distance is shortest on tasks with abundant public training data and clear evaluation signals. It is longest on:

  • long-horizon agentic tasks, where post-training on proprietary interaction data matters;
  • multimodal understanding at high resolution and across video; and
  • reliability at the tail: the rare, expensive failures that enterprise buyers care about most.

Those three categories are exactly where the closed labs are concentrating their effort, which is not a coincidence.

Why the gap exists

A frontier training run costs hundreds of millions of dollars in compute alone, before the research staff and the data. Somebody has to expect a return. For the closed labs the return comes from exclusive access: API revenue, consumer subscriptions and enterprise contracts that only exist because nobody else can serve the same model. Releasing the weights destroys that return.

Open-weight releases therefore come from organisations whose return does not depend on exclusivity:

FunderWhy release weightsWhat they tend to release
Hardware vendorsEvery open model sells more chipsModels optimised for their own silicon
Cloud platformsHosting is the product; the model is a loss leaderBroad general-purpose models
Labs with a services businessConsulting, fine-tuning, on-prem deploymentsStrong base models with permissive licences
National or strategic programmesSovereignty, talent, influenceMultilingual models and large releases with unusual timing
Labs seeking distributionMindshare that converts to a later closed tierExcellent small and mid-size models

Five columns of open-weight funders: hardware vendors, cloud platforms, services labs, national programmes and distribution seekers, each with why they release weights and what they tend to release

Figure 1: Open weights come from organisations whose return does not depend on exclusive access.

Each of these funders has a reason to stop at “good enough to be useful and widely adopted”. None of them has a reason to spend a further billion dollars to be first by three months, because being first by three months is only worth a billion dollars if you can charge for exclusive access. That is the structural reason the frontier stays closed.

Two lanes: a frontier lab funds a run, expects a return from exclusive access and keeps weights closed; a chip, cloud or state funder earns its return elsewhere, wants adoption and ships weights

Figure 2: The gap at the frontier is a funding question, not a technical one.

Where that leaves builders

For most production workloads the decision is no longer binary. The practical questions are:

  1. Is the task within reach of the current open frontier? For classification, extraction, summarisation, RAG and most coding assistance, yes, and has been for a while. Re-benchmark quarterly and you will keep finding that the answer expanded.
  2. Where will you serve it? Self-hosting makes sense at high, steady volume or under data-residency constraints. Otherwise a hosted open-model provider is usually cheaper than owning GPUs.
  3. What does the licence actually allow? “Open” spans everything from Apache 2.0 to licences with user thresholds, field-of-use restrictions and attribution requirements. Read it.
  4. Who maintains the model? Open releases are snapshots. Security patches, tokenizer fixes and safety updates depend on the releasing lab continuing to care.

The frontier will stay closed for the tasks that need it. The rest of the industry is increasingly running on weights anyone can download, and that is the more consequential trend.

Frequently asked questions

What is the difference between open-weight and open-source models?

Open-weight models publish their trained parameters so anyone can run and fine-tune them, but often under licences with usage restrictions and without the training data or code. Open-source, strictly used, would require an OSI-approved licence and typically the training recipe too. Most "open" models today are open-weight.

Will open-weight models catch up completely?

On most measurable tasks they are already close, and for many production use cases the difference is immaterial. At the absolute frontier the gap is likely to persist because releasing weights removes the exclusivity that funds the largest training runs.

All models →