🔥 21 cent 📚 Learn

LEARN

The cost of running open models at scale

Companies can download an open-weight model without paying for access to an API, but large models still require substantial computing infrastructure when deployed at scale.

Kimi K3 contains 2.8 trillion total parameters and 104 billion activated parameters, according to Moonshot. Its mixture-of-experts architecture includes 896 experts, with 16 selected for each token.

The model’s size still places substantial hardware requirements on operators. Moonshot temporarily stopped accepting new Kimi K3 subscriptions in July after saying usage had placed pressure on its available GPUs, while Reuters reported that relatively few users were expected to self-host a model of that scale because of the infrastructure required.

Alibaba is using a similar architectural approach with Qwen3.8-Max. The model contains about 2.4 trillion parameters but activates around 95 billion parameters for each request, according to Reuters.

Moonshot says its mixture-of-experts design improves scaling efficiency by activating only a subset of the model’s experts for each token, rather than the full model.

Cloud providers can charge for hosting and inference, while AI infrastructure companies can generate revenue from deployment and optimisation services.

Dan Fu, vice president of kernels at Together AI, said companies providing AI services can differentiate their offerings through areas such as more efficient token use and deployment optimisation.

“At the application layer, there’s value out there for how you use it, how you actually get the models and the tokens to do something useful,” Fu said.

Model development presents a separate cost challenge. Research involving Epoch AI and Stanford researchers estimated that the cost of the most compute-intensive training runs had risen by about 2.4 times a year since 2016, while Stanford’s 2025 AI Index found that the price of accessing models at a given capability level had fallen sharply.

At published API prices at the time of release, Kimi K3 was priced at about one-third of Anthropic’s Fable model based on listed input and output token rates. Pricing is only one part of the deployment cost, particularly for companies running models on dedicated infrastructure or handling high volumes of requests.

These costs sit alongside the licensing arrangements being tested by model developers. Alibaba already charges developers that access Qwen through Alibaba Cloud. The proposed arrangement would also allow it to collect revenue from some companies deploying Qwen independently on their own infrastructure or through third-party services.

Moonshot has already attached commercial conditions to Kimi K3 while keeping its model weights available for download. DigitalOcean and Chinasoft International have both disclosed commercial arrangements with Moonshot, although the financial terms have not been made public.

The commercial arrangements are developing alongside wider tensions between China and the US over AI technology. The White House has accused Moonshot of using technology taken from Anthropic while developing its models, an allegation Chinese officials have rejected.

Interest in releasing models with downloadable weights is not limited to Chinese developers. Thinking Machines Lab, the San Francisco AI company founded by former OpenAI Chief Technology Officer Mira Murati, released its first open-source model last month.

Lin Qiao, chief executive and co-founder of Fireworks AI, said there was no fundamental technical barrier preventing US developers from releasing more capable open-source models. Fireworks AI works with models from developers including Moonshot, although Qiao declined to discuss its commercial arrangements.

Alibaba has not publicly announced the final licence for its next Qwen model or the revenue-sharing percentage it plans to seek from large commercial users.

← BACK TO 21 CENT