The performance advantage of the most advanced proprietary artificial intelligence models is shrinking, raising questions about whether businesses need to pay a substantial premium for frontier AI systems. A Mozilla report finds that leading open-weights models are now only about 4.4 months behind their closed counterparts while offering significantly lower operating costs.
For Canadian businesses and technology teams assessing AI spending, the findings suggest open models could increasingly handle routine workloads, while more expensive closed systems may remain worthwhile for a narrower range of demanding tasks.
Open Models Gain Ground on Frontier AI Systems
Mozilla’s State of Open Source AI report argues that most organizations should consider open models as the default for much of their AI work.
One leading open-weights model, Moonshot AI’s Kimi K3, achieved a composite performance score on the Artificial Analysis Intelligence Index only three points behind Anthropic’s closed Fable 5 frontier model. Kimi K3 costs about 30 per cent as much.
“[A Closed model] earns its premium in a few places: expert professional work, high-intensity retrieval, and long context,” Mozilla chief technology officer Raffi Krikorian said. “We see the decision to pay for closed [models] as workload-specific rather than organization-specific.”
Open-weights AI models allow users to download key model components and operate them on their own infrastructure. However, they are not necessarily fully open source because developers may still withhold information including training data, data pipelines and training code.
By comparison, companies such as Anthropic and OpenAI generally keep their frontier models proprietary and charge customers for access.
Closed Models Still Offer Convenience
Organizations continue paying for proprietary systems partly because they are easier to deploy and often include compliance tools, technical support and clearer accountability.
Many companies also lack the specialized staff and computing infrastructure required to operate open models efficiently.
However, improving performance and lower costs are making open alternatives increasingly attractive. Delivery company DoorDash, for example, has used Kimi for routine work while reserving Fable for more difficult assignments that would otherwise require significant time from human experts.
Performance Gap Falls to About Four Months
Research organization METR measures AI capabilities partly through a model’s “time horizon.” This represents the length of a task, based on how long a human expert would require to complete it, that an AI model can handle with a reliable 50 per cent success rate.
According to Mozilla’s analysis, the strongest closed model can currently complete a task approximately 1.7 times longer than the longest assignment the best open model can reliably finish.
“If the open frontier can handle a seven-hour job, the closed frontier can handle a 12-hour one,” Krikorian said. “In four months, the open model handles the 12-hour job, and the closed one handles something around 20.”
The most significant difference currently appears in tasks requiring roughly eight to 12 hours of expert human work. Closed frontier models may be able to handle these assignments while existing open alternatives cannot yet complete them reliably.
Tasks requiring less than eight hours could increasingly be assigned to cheaper open models. Tasks longer than 12 hours, meanwhile, generally remain beyond the reliable capabilities of either category.
Neutral Testing Highlights Cost Advantage
Comparing AI systems is complicated by the software surrounding each model. Frontier models often include proprietary “harnesses” that provide access to tools, memory and other features needed for agent-like tasks.
Benchmarking company Vals AI sought to reduce this advantage by evaluating open and closed systems using the same neutral harness.
In the Terminal-Bench 2.1 evaluation, Chinese developer Z.ai’s open-weights GLM 5.2 scored within one point of Anthropic’s Claude Opus 4.7 and 4.8 while costing roughly five times less per completed task.
The results suggest organizations may effectively be paying around five times more for a performance advantage that open models could close within approximately four months, particularly for workloads in the eight-to-12-hour range.
“Pay when that head start is worth it, something like a deadline that lands before the open frontier catches up would be here,” Krikorian said. “Routine work you’ll still be doing next quarter is not, because you’ll be able to do it for a fifth of the cost soon, and the model won’t be the bottleneck anyway.”
Chinese Developers Lead Open-Weights AI
Open-weights models are also gaining substantial usage. Mozilla found that eight of the 10 most-used models by token volume on AI marketplace OpenRouter offered open weights.
Revenue, however, remains heavily concentrated among closed-model providers. Research by Frank Nagle and Daniel Yue for the Linux Foundation found that open models accounted for just four per cent of overall revenue during the period examined, compared with 96 per cent for closed systems.
“In the last year, open models have exploded, so we do expect revenue to have shifted,” Krikorian said.
Geographic concentration is another concern. Many of the strongest open models currently come from China, while leading closed frontier systems are primarily developed by U.S. companies.
“The uncomfortable truth is that the plural ecosystem around open-weight AI is largely funded by Chinese capital right now,” Krikorian said. “That’s a plurality and a concentration at the same time.”
Push for a Broader Open AI Ecosystem
Krikorian argues that organizations in the United States and Europe should compete more aggressively in open AI development to reduce the risk of a single country establishing widely adopted global standards.
He pointed to Linux and other open-source infrastructure as possible examples, where companies, foundations and public institutions collectively supported technologies that benefited the wider ecosystem.
Public computing programs, neutral foundations, private companies and philanthropic organizations could play similar roles in AI development, particularly by supporting transparent evaluation and auditing systems.
Krikorian also argued that the industry should eventually move beyond releasing model weights and provide greater transparency about how models are trained and evaluated.
“It’s hard to fully trust a model with decisions if you can’t tell how it was trained or what it was evaluated against,” he said.
As open models continue improving, the economic case for expensive frontier AI is becoming increasingly dependent on individual workloads. For Canadian organizations managing growing technology costs, the key question may no longer be which AI model is the most powerful, but whether a temporary performance advantage is valuable enough to justify paying several times more.

Nolan Fraser is a contributor at Angperyodiko.ca, covering a wide range of topics including news, politics, business, technology, sports, entertainment, and lifestyle. He focuses on delivering clear, accurate, and accessible reporting that helps readers stay informed about current events and developments that matter. With an emphasis on useful information and balanced storytelling, Nolan aims to provide timely coverage and relevant insights for audiences across Canada and beyond.