AiPhreaks ← Back to News Feed

Beyond VLAs: How World Action Models Reshape Robot Manipulation

By Jakub Antkiewicz

2026-08-05T10:33:25Z

A Foundational Shift in Robot Learning

A notable shift is underway in robot policy development, moving from established Vision-Language-Action (VLA) models to a new architecture known as World Action Models (WAMs). NVIDIA is promoting this approach with its recently opened Cosmos 3 foundation model, which is designed to provide robots with a stronger grasp of physical dynamics. This matters because WAMs enable policies to generalize to new tasks, objects, and robot hardware with substantially less task-specific data, addressing a core challenge in creating versatile manipulation systems.

From Semantics to Physics

The limitation of conventional VLAs is that their backbones are trained to describe the world, not to predict how it will physically evolve. A WAM, by contrast, is built upon a video world model, giving it an innate physics prior. Instead of just mapping language instructions to actions, it learns cause and effect from vast datasets of interaction. NVIDIA's open Cosmos 3 model serves as a foundation for building these policies and is distinguished by its architecture and training data.

  • Architecture: Built on a Mixture-of-Transformers (MoT) that handles multimodal inputs.
  • Training Data: A vast dataset including approximately 767M images, 348M videos of real-world dynamics, and 8M action samples.
  • Deployment Tiers: The model is available in multiple sizes, including a 4B parameter Cosmos Edge model for on-device inference on NVIDIA Jetson and a 16B Cosmos Nano model for workstation serving.
  • Functionality: The model can jointly predict robot actions and the resulting future video frames, essentially imagining the outcome as it acts.

Impact on the Robotics Ecosystem

By open-sourcing the Cosmos 3 model, datasets, and training recipes under a commercial-use license, NVIDIA aims to lower the barrier for developing sophisticated robot policies. The practical benefit for engineering teams is a significant reduction in the data and time needed to adapt a policy to a new robot embodiment or task. This approach standardizes the starting point for development, allowing teams to build upon a shared foundation of physical understanding rather than pretraining a new model from scratch for each unique application, from Franka Panda arms to dual-arm setups.

Strategic Takeaway: NVIDIA's promotion of WAMs over VLAs marks a strategic pivot from teaching robots to mimic actions to enabling them to understand physical consequences. By focusing on a generalizable physics prior, this approach reduces the data dependency and engineering costs associated with adapting policies to new hardware, potentially accelerating the deployment of physically competent robots in unstructured environments.
End of Transmission
Scan All Nodes Access Archive