Chinese LLMs' Cost Advantage Keeps Widening, UBS Securities Says(Yicai) July 27 -- Chinese large language models keep winning market share on the back of a structural cost advantage rather than price subsidies, with this edge continuing to widen, according to an internet analyst at UBS Securities China.
The closely watched release of Kimi K3 by Beijing-based startup Moonshot AI on July 16 reinforced this view, Xiong Wei said in a recent media briefing. Leading Chinese artificial intelligence models cost about a tenth to train compared with their top overseas peers, while their average application programming interface pricing is around 10 percent to 20 percent of rivals, she added, citing data compiled by UBS.
Kimi K3 is priced at the high end among Chinese models, yet it still costs only about 30 percent of comparable overseas frontier models, Xiong noted. Even so, domestic model providers' API businesses maintain gross margins of 20 percent to 40 percent, which is lower than the latest margins reported by their US peers, but still healthy, she stressed.
Part of the training-cost gap stems from the generally smaller parameter counts of Chinese models, but their bigger driver is innovation in algorithms and engineering, according to Xiong. Compared with US developers that prioritize absolute performance gains and are willing to pour research spending into frontier exploration, Chinese companies focus on balancing performance with cost efficiency, a reflection of China's relatively constrained compute supply and developers' broader push for AI accessibility, she pointed out.
China also has an active open-source ecosystem, in which leading labs publish research alongside each new model generation, turning this into another factor that helps technological progress spread quickly across teams and creates a collective momentum, she said, adding that this is in contrast with the largely closed, siloed approach of leading US models.
Kimi K3's release should strengthen rather than undermine market confidence in China's open-source ecosystem, Xiong stressed, arguing that a strong new model is a net positive for the broader ecosystem rather than simply a share grab from other providers.
On inference costs, the advantage of Chinese models can be broken down into three layers, she pointed out. The first is model architecture, where under mixture-of-experts designs, mainstream models typically activate a single-digit to 10 percent share of parameters per inference task, versus an estimated 15 percent to 30 percent for US models, directly affecting the computing power required for the same task, she said.
The second is engineering and scheduling efficiency, with industry-wide graphics processing unit utilization in inference averaging around 40 percent to 50 percent based on UBS research, while leading Chinese players have pushed utilization above 70 percent through scheduling optimization, she noted. The third is infrastructure expenses, with cheaper electricity and data-center costs in China providing an additional structural saving, one that could widen further over the long term due to more inference workloads shifting to domestic chips, she added.
Demand-side trends are also moving in Chinese providers' favor, Xiong pointed out. Earlier this year, the industry embraced so-called "token maxxing," encouraging users to consume as many tokens as possible, she said, but noted that as AI bills have piled up, companies are shifting toward budget management that prioritizes return on investment.
As task requirements become more stratified, the cost-effectiveness of Chinese models is becoming a more prominent factor in enterprise model selection, she said, giving examples of overseas companies that have begun shifting parts of their workflows to Chinese open-source models.
For the second half of this year, there are three areas worth watching, Xiong said. The data flywheel effect created by the industry's push into coding is likely to help leading developers consolidate their edge; coding-agent products are expanding from software engineers to a broader range of white-collar knowledge workers, an area where Chinese firms' multimodal and video-generation capabilities also show relative strength; and the pricing-power gap between top-tier models and others is likely to become more pronounced because of demand becoming more stratified, she pointed out.
The main constraint on revenue growth for Chinese model providers is no longer demand but compute supply, Xiong stressed. Most providers do not have enough computing power dedicated to inference to meet actual demand, so continued progress in the number and performance of domestic chips will determine how much of that demand they are ultimately able to convert into income, she added.
Editor: Martin Kadiev
