Bigger Models or Multimodal? Experts Debate AI's Future Path
Lv Qian
DATE:  a day ago
/ SOURCE:  Yicai
Bigger Models or Multimodal? Experts Debate AI's Future Path Bigger Models or Multimodal? Experts Debate AI's Future Path

(Yicai) July 20 -- The development trend of artificial intelligence models has turned into one of the hottest topics at the World AI Conference in Shanghai, with some industry experts pointing to larger-scale parameters as the standard for next-generation models, while others see multimodal capabilities as the future direction.

Parameters

The AI model development trend is toward the two trillion parameter size, a representative from Shanghai-based AI startup MiniMax Group told Yicai at the four-day WAIC, which ends today. US' Anthropic has already validated the path, the person noted.

Anthropic is the world's most valued AI startup, best known for its Claude large language model series. The Claude Fable 5, released last month, is estimated to be in the five trillion to 10 trillion parameter range.

MiniMax's next-gen model, codenamed M3 Pro, is also rumored to have over two trillion parameters, focusing on enhancing complex inference and multi-step agent tasks.

Beijing-based startup Moonshot AI also attracted global attention after launching its next-gen Kimi K3 with 2.8 trillion parameters on the eve of the WAIC. Gavin Baker, founder of US tech investment firm Atreides Management, said on X that the model could have negative impacts on industry-leading companies like Anthropic and OpenAI.

In addition, iFlytek is developing trillion-level parameter models, a representative from the Chinese AI voice tech developer told Yicai. The industry has gradually shifted into two directions: models with hundreds of billions of parameters aimed at handling daily and simple tasks and models with trillions of parameters focused on complex reasoning, agent automation, and advanced tasks, the person pointed out.

Capabilities

As AI moves from the virtual to the real world, multimodal capabilities will become key to model evolution, several experts said at a WAIC roundtable forum.

The intelligence evolution path depends on multimodal interaction, because a single language model only compresses static knowledge from text and cannot replicate the spatiotemporal relationships, causal logic, and behavioral interactions of the physical world, Zhang Xiangyu, chief scientist at Shanghai-based AI startup StepFun.

Although language models can undertake tasks in digital scenarios, only native multimodal capabilities can support the emergence of intelligence in manufacturing, embodied intelligence, and other physical scenarios, noted Liu Ziwei, associate professor at Nanyang Technological University in Singapore.

The biggest shortcoming of AI models is the lack of contextual understanding of real scenarios, said Qiu Xipeng, a professor at Fudan University. Future general intelligence will create a new form of foundational models + multimodal interaction infrastructure, relying on engineering systems to compensate for real-world interaction cognitive shortcomings, Qiu added.

Editors: Dou Shicong, Martin Kadiev

Follow Yicai Global on
Keywords:   WAIC,LLM,Multimodal