In the first seven months of 2026, Chinese startups in the world model space raised a combined 66.6 billion yuan (approximately $9.1 billion)—a more than fivefold increase compared to the same period in 2025. Beijing alone accounts for 46% of the country's world model startups, cementing the capital as the global epicenter of this emerging AI frontier. From venture capital firms to state-backed funds, from tech giants to university spinouts, the message is unmistakable: world models are the next great bet in artificial intelligence, and China is determined to lead. But what exactly are world models, why are they attracting such staggering investment, and can China translate its financial firepower into lasting technological dominance?

¥66.6B
World Model Financing (Jan–Jul 2026)
5×+
Year-over-Year Growth
46%
Beijing's Share of Startups
3
Major Technical Schools
100+
Active World Model Startups
$9.1B
USD Equivalent

What Are World Models, and Why Do They Matter?

To understand the investment frenzy, one must first understand what a world model is. Unlike large language models (LLMs) that process and generate text, world models are AI systems designed to build internal representations of physical environments—simulating how objects move, interact, and change over time. Think of a world model as an AI's internal "imagination engine": it can predict what happens next in a scene, reason about three-dimensional spatial relationships, and generate realistic video sequences that obey the laws of physics.

The concept has deep roots in AI research. In 2018, DeepMind's David Ha and Jürgen Schmidhuber published a seminal paper describing world models as systems that learn compressed representations of environments and use them to plan actions. The idea was that a truly intelligent agent needs an internal "mental model" of the world. Seven years later, the convergence of transformer architectures, diffusion models, and massive compute resources has turned this theoretical framework into a commercially viable technology.

Why does this matter? Because world models are the missing link between today's AI—which excels at pattern recognition in static data—and tomorrow's AI, which must understand dynamic physical reality. Applications span autonomous driving (predicting how traffic scenarios unfold), robotics (simulating manipulation tasks before execution), video generation (creating physically consistent films), gaming (building interactive worlds that respond realistically), and scientific simulation (modeling molecular interactions or climate dynamics). In short, world models are the pathway to AI that understands the world the way humans do: through cause and effect, space and time.

Why World Models Are Different from LLMs

LLMs process sequences of tokens. World models process sequences of sensory observations—video frames, depth maps, LiDAR point clouds, and action trajectories. This fundamental difference means world models must learn not just statistical patterns in language, but the causal structure of physical reality. A language model can tell you that "if you drop a glass, it will break." A world model can simulate the glass falling, predict the exact moment of impact, and generate a video of the shattering process consistent with the material properties of glass, the angle of impact, and the surface it lands on.

The Three Technical Schools Competing for Dominance

China's world model ecosystem is not a monolith. Three distinct technical approaches have emerged, each backed by different research traditions and commercial visions. The competition between these schools is shaping the trajectory of the entire field.

1. The Pixel Generation School

The pixel generation approach treats world models primarily as video generation engines. These systems take text prompts or reference images and produce high-fidelity video sequences. The core insight is that a model capable of generating realistic, temporally consistent video must have internalized the physics of the world—how light reflects off surfaces, how fluids flow, how objects deform under force. The dominant architectures are diffusion transformers (DiT) and their variants, trained on billions of video-text pairs.

This school is the most commercially visible. Its products are directly marketable: text-to-video generation tools, AI filmmaking assistants, and content creation platforms. Companies in this category have attracted the largest funding rounds because their path to revenue is the clearest. However, critics argue that video generation alone does not constitute a true world model—generating plausible-looking video is not the same as understanding the underlying causal structure.

2. The Action-Conditioned School

The action-conditioned approach, sometimes called the "interactive world model" school, focuses on models that respond to agent actions. These systems take as input not just sensory data but also an action (e.g., "move forward," "turn left," "pick up object") and predict the resulting sensory observation. This is the architecture closest to the original Ha and Schmidhuber vision: a model that an agent can use to plan by simulating the consequences of its actions before executing them in the real world.

This approach is particularly relevant for robotics and autonomous driving. A robot equipped with an action-conditioned world model can "imagine" thousands of possible manipulation sequences, evaluate the outcomes, and select the one most likely to succeed—all before touching a physical object. The challenge is that action-conditioned models require training data that includes both sensory observations and action labels, which is far scarcer than raw video data.

3. The Latent Space School

The latent space school, sometimes aligned with the "world model as scientific instrument" philosophy, focuses on learning compressed representations of physical dynamics. Rather than generating pixels, these models learn to predict future states in an abstract latent space—a lower-dimensional representation that captures the essential physics of a system while discarding irrelevant details. This approach draws heavily from the neural ODE (ordinary differential equation) and neural operator traditions in scientific machine learning.

Proponents argue that latent space world models are more scientifically rigorous: they can be validated against physical laws, they generalize better to unseen scenarios, and they provide interpretable representations that scientists can analyze. The trade-off is that they are harder to commercialize in the short term. A latent space model that accurately predicts fluid dynamics is valuable for aerospace engineering but does not generate the kind of eye-catching demos that attract venture capital. Still, this school has deep connections to China's national laboratories and defense research ecosystem, giving it a different kind of sustainability.

The Schools at a Glance

Pixel Generation: Text/video → video output. Best demos, clearest revenue path. Question: is it really a world model or just a very good video generator?
Action-Conditioned: Observation + action → next observation. Most relevant to robotics and autonomy. Scarcer training data.
Latent Space: Learn compressed dynamics. Most scientifically rigorous. Hardest to commercialize. Strongest ties to national labs.

The Four Stars of China's World Model Race

Among the 100-plus world model startups in China, four companies have emerged as the clear frontrunners—each representing a different approach to the technology and a different bet on its commercial future.

Jijia Vision (极佳视界)

Founded in 2023 by a team of former Alibaba DAMO Academy and Tsinghua University researchers, Jijia Vision has positioned itself as the leader of the pixel generation school. The company's flagship product, a text-to-video and image-to-video generation platform, has been adopted by over 200 corporate clients in advertising, film production, and e-commerce. Jijia Vision closed a Series B round in March 2026 at a reported valuation of $2.8 billion, with investors including IDG Capital, Sequoia China, and a subsidiary of the National Integrated Circuit Industry Investment Fund. The company's technical approach combines a proprietary diffusion transformer with a physics-aware training objective that penalizes physically impossible outputs—a hybrid strategy that bridges the pixel generation and latent space schools.

Manifold Space (流形空间)

Manifold Space, a Shanghai-based startup spun out of Shanghai Jiao Tong University's AI Institute, is the standard-bearer of the action-conditioned school. The company has built a world model specifically designed for industrial robotics, enabling robotic arms to simulate grasp and manipulation sequences before execution. In June 2026, Manifold Space announced a partnership with four of China's largest electronics manufacturers to deploy its world model in production lines, reducing robot programming time by an average of 70%. The company raised 3.2 billion yuan in a Series A+ round led by Hillhouse Capital and BYD's corporate venture arm, reflecting the strategic importance of industrial automation to China's manufacturing sector.

Inverse Matrix (逆矩阵)

Inverse Matrix occupies the frontier where the latent space and action-conditioned schools converge. The company, headquartered in Beijing's Haidian district—the same neighborhood that houses 46% of China's world model startups—has developed a world model architecture that operates in a compressed latent space while still supporting action-conditioned rollouts. Its technology is being used by China's State Key Laboratory of Autonomous Systems to simulate autonomous vehicle scenarios at scale, testing edge cases that would be dangerous or impractical to reproduce physically. Inverse Matrix has raised approximately 5.8 billion yuan across two funding rounds in 2025–2026, with a significant portion coming from the Beijing Municipal AI Industry Investment Fund, a government vehicle explicitly created to support world model development.

HiDream.ai

HiDream.ai is the wildcard. Founded by a team with experience at both Google DeepMind and ByteDance, the company has pursued a deliberately multimodal approach, building a world model that jointly handles text, images, video, and 3D geometry. HiDream's platform generates 3D-consistent scenes from text descriptions—a capability that has attracted attention from the gaming and architectural visualization industries. The company closed a $420 million funding round in April 2026, co-led by Tencent and GGV Capital. HiDream's ambition is to build a "foundational world simulator" that can serve as the physics engine for any AI application that needs to reason about physical space.

Jijia Vision

Location: Hangzhou / Beijing

Focus: Text-to-video generation, physics-aware diffusion

Valuation: ~$2.8B (Series B, Mar 2026)

Key Investors: IDG Capital, Sequoia China, National IC Fund

Manifold Space

Location: Shanghai

Focus: Industrial robotics world models, action-conditioned simulation

Funding: ¥3.2B (Series A+, Jun 2026)

Key Investors: Hillhouse Capital, BYD Ventures

Inverse Matrix

Location: Beijing (Haidian)

Focus: Latent space models, autonomous driving simulation

Funding: ¥5.8B (2025–2026 total)

Key Investors: Beijing Municipal AI Fund

HiDream.ai

Location: Shenzhen / Beijing

Focus: Multimodal world simulation, 3D-consistent generation

Funding: $420M (Apr 2026)

Key Investors: Tencent, GGV Capital

Policy Drivers: Why Beijing Is All-In

The 66.6 billion yuan flowing into world models is not purely market-driven. Behind the investment numbers lies a coordinated policy push that has accelerated the formation of China's world model ecosystem.

In January 2026, China's State Council released the "Next-Generation Artificial Intelligence Development Plan (2026–2030)," which explicitly identified world models as one of five "strategic AI technologies" deserving priority funding. The document called for China to achieve "world-leading capabilities" in world model research by 2028 and to establish a "complete industrial ecosystem" by 2030. The plan was accompanied by concrete mechanisms: a dedicated 50 billion yuan national fund for world model research, preferential tax treatment for world model startups, and streamlined approval processes for grant applications.

At the municipal level, Beijing has been the most aggressive. The city's Haidian district, already home to China's densest cluster of AI companies, launched a "World Model Innovation Hub" in February 2026, offering subsidized office space, compute credits on state-backed cloud platforms, and matching funds for venture capital investments. The 46% figure—Beijing's share of national world model startups—is not an accident; it is the result of deliberate industrial policy.

Shanghai, Shenzhen, and Hangzhou have followed with their own initiatives, creating a competitive dynamic among cities that mirrors the earlier race to attract AI chip design companies. The competitive federalism of China's tech policy means that world model startups can often negotiate favorable terms by playing cities against each other, further accelerating the flow of capital into the sector.

Global Competition: Where China Stands

China is not alone in pursuing world models. The technology has become a global priority, with major investments from the United States, Europe, and other AI research hubs. Understanding the competitive landscape requires a clear-eyed comparison.

DimensionChinaUnited StatesEurope
Total Funding (2026 YTD)¥66.6B (~$9.1B)~$4.2B~€1.8B
Number of Dedicated Startups100+~40~25
Government Funding ProgramsNational + Municipal (Dedicated)DARPA, NSF (General AI)Horizon Europe (General AI)
Leading Research LabsTsinghua, SJTU, Peking U, CASMIT, Stanford, Berkeley, DeepMind USOxford, ETH Zurich, INRIA
Key Corporate PlayersAlibaba, Tencent, ByteDance, BaiduGoogle, Meta, OpenAI, NVIDIAMistral, Aleph Alpha
Primary Commercial FocusContent generation, robotics, autonomous drivingGaming, scientific simulation, foundation modelsClimate modeling, digital twins

The funding disparity is striking but should be interpreted carefully. Chinese venture capital data often includes commitments from state-backed funds that may be disbursed over multiple years, making direct comparisons with US venture capital figures difficult. Moreover, the US ecosystem benefits from deeper integration between world model research and the existing AI infrastructure stack—NVIDIA's GPUs, Google's TPUs, and the major cloud platforms all provide world model researchers with tools that have no direct Chinese equivalent.

On the research front, the picture is more balanced. Chinese universities and corporate labs have produced world model papers that rank among the most cited in the field. Tsinghua University's "WorldDreamer" architecture and Shanghai AI Lab's "DriveDreamer" for autonomous driving simulation have been widely adopted as baselines. However, the foundational breakthroughs—the transformer architecture, diffusion models, and the original world model formulation—remain overwhelmingly Western in origin. China's strength is in applied innovation and scaled deployment, not in the kind of paradigm-shifting theoretical advances that redefine a field.

Risks and Challenges: What Could Go Wrong

For all the enthusiasm, the world model sector faces significant risks that could derail the investment thesis.

The Valuation Question

The most immediate concern is valuation. When a sector sees 5× year-over-year funding growth, the question of whether capital is being allocated efficiently becomes unavoidable. Several of the 100-plus world model startups in China have valuations exceeding $500 million despite having no commercial product, no revenue, and in some cases, no publicly demonstrated technology beyond research papers. The pattern is reminiscent of the autonomous driving investment wave of 2017–2019, which saw dozens of Chinese self-driving startups raise billions of dollars before consolidation wiped out the majority of them. A similar shakeout in world models is not just possible—it is likely.

Technical Hurdles

World models face fundamental technical challenges that are not yet solved. Training a world model that can simulate a scene with high fidelity for more than a few seconds—maintaining temporal consistency, respecting physical constraints, and handling long-range dependencies—remains an open research problem. Most current systems produce convincing short clips but degrade rapidly when asked to simulate extended sequences. The computational cost is also enormous: training a state-of-the-art video world model can require tens of thousands of GPU-hours, putting it out of reach for all but the best-funded labs.

Data Scarcity for Action-Conditioned Models

While the pixel generation school can train on the vast supply of publicly available video data, the action-conditioned school faces a data bottleneck. Training a world model that can predict the consequences of robot actions requires paired data—video of robot actions with corresponding action labels—which is expensive to collect and difficult to scale. This bottleneck may limit the near-term commercial viability of the approach that is most directly relevant to industrial applications.

Geopolitical Technology Restrictions

US export controls on advanced GPUs continue to constrain Chinese AI companies. While world model training can be done on domestically produced AI chips from companies like Huawei (Ascend series) and Biren Technology, the performance gap with NVIDIA's latest offerings remains significant. If export controls tighten further—a possibility under any US administration given the bipartisan consensus on technology competition with China—the computational foundation of China's world model ambitions could be undermined.

The Consolidation Thesis

Industry analysts broadly expect that the 100-plus world model startups in China will consolidate to fewer than 20 within three years. The winners will likely be those that combine strong technical teams with clear commercial applications and deep relationships with either state-backed funds or major corporate partners. The losers will be those that raised money on the strength of demos and hype without a sustainable path to revenue. The question for investors is not whether the technology is important—it is—but whether they are backing the right horse in a race where most runners will not finish.

What World Models Mean for the Broader AI Landscape

Beyond the investment numbers, the rise of world models signals a structural shift in how the AI industry thinks about progress. For the past five years, the dominant narrative has been about scaling—bigger models, more data, more compute. World models represent a different kind of ambition: not just scaling pattern recognition, but building systems that internalize the causal structure of reality.

If world models succeed, the implications extend far beyond the companies building them. An AI system with a robust internal model of physical reality could transform industries that have been resistant to earlier waves of AI automation: construction, where robots need to understand how materials behave; healthcare, where surgical robots need to predict tissue deformation; and scientific research, where simulations of molecular dynamics could accelerate drug discovery. The economic value of a functional world model is, by any reasonable estimate, measured in trillions of dollars over the coming decades.

China's bet on world models is, in this sense, a bet on the next phase of AI itself—a phase where language models are commoditized and the frontier shifts to systems that can reason about the physical world. Whether that bet pays off depends on factors that no amount of capital can guarantee: the resolution of fundamental research questions, the emergence of sustainable business models, and the ability of Chinese companies to compete on innovation rather than just scale. The 66.6 billion yuan is a down payment on a future that has not yet been built. The world is watching to see whether it was wisely placed.