
How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Machine Learning Street Talk (MLST)
NVIDIA’s Cosmos 3 model functions as a versatile "omni-model" that integrates text, video, audio, and action to advance physical AI and robotics. By serving as a neural simulator, it enables developers to verify robot policies without relying solely on costly real-world testing, effectively bridging the sim-to-real gap. The architecture utilizes a vision-language model foundation combined with a bidirectional diffusion generator to capture complex correlations between visual observations and actions. This framework supports diverse embodiments, allowing for better generalization across tasks like navigation and complex manipulation. With three distinct versions—Super, Nano, and Edge—the platform provides scalable solutions for both high-fidelity research and on-device deployment. By open-sourcing these models, data, and training recipes, the initiative aims to accelerate development velocity and safety standards for autonomous systems and humanoid robotics.
Sign in to continue reading, translating and more.
Open full episode in Podwise