A new approach called “Code-as-World” is turning real-world videos directly into executable physics programs. This agentic loop uses AI to analyze footage and automatically generate MuJoCo simulation code, marking a significant step toward bridging the gap between perception and physical simulation. For robotics and AI training, this could mean generating vast, realistic training environments from simple video clips.

What Happened

The method works by feeding a video into an AI agent that deconstructs the scene, identifies objects, materials, and interactions, and then writes a complete MuJoCo program that simulates that exact physical scenario. The system can handle complex dynamics like collisions, deformations, and fluid-like behavior by rewriting the video’s content into code that runs in the MuJoCo physics engine.

This is not just about generating a 3D model — it’s about creating an executable, interactive simulation that obeys physics laws. The agent uses a loop: it generates the program, runs it, compares the output to the original video, and iteratively refines the code until the simulation matches the observed reality.

The implications are enormous for robotics. Instead of manually designing simulation environments for reinforcement learning, researchers could simply show a robot a video of a task, and the system would automatically create the training ground. The same goes for game development, autonomous vehicle training, and industrial automation.

Read the full announcement →

My Take

This is the kind of progress that actually matters for embodied AI. The bottleneck in robotics reinforcement learning has always been the simulation-to-reality gap and the manual effort required to build realistic training environments. Code-as-World addresses both by treating video as a direct input to simulation generation.

The key innovation here is the iterative refinement loop. The system doesn’t just guess — it plays the simulation back, checks it against the source video, and fixes its mistakes. This feels like a genuine step toward self-supervised simulation building, which could dramatically accelerate how quickly robots learn physical manipulation and navigation tasks.

What to Watch

  • Whether this technique can handle highly complex, multi-object scenes with occlusions and transparent materials
  • If the generated MuJoCo programs can be directly used for training reinforcement learning policies
  • Potential integration with large language models to enable natural language commands like “simulate a video of someone pouring water into a cup”