Two YAM arms and the block-in-bin cell, simulated in MuJoCo in this tab. A small learned renderer turns the simulator's view into what the overhead camera would see, live.
Physicsloading MuJoCo…
Rendererloading model…
Vieworbit · drag to turn, scroll to zoom
How this works
The cell is MuJoCo. The arms, table, mat, block and bin are the compiled twin model the training data was rendered from. In Drive, MuJoCo itself runs here (WebAssembly, 1 ms steps): you set the arms' motor targets, and gravity, contacts and the grasp do the rest.
The table is real. Its texture is the overhead camera's photo of the empty table, assembled from hundreds of frames with the arms masked out.
The cameras are the rig's. Overhead is the registered overhead camera; the wrist views ride on the grippers at the mounts calibrated from the wrist footage.
Rendered is a learned model (v1, in this tab) turning MuJoCo's overhead view into a real-looking frame. Replay shows held-out recordings next to the real camera.
Physics is not tuned to the real arms' dynamics. The v1 renderer only knows this cell, and poses far from the recordings render poorly.
Sourcea held-out recording, played through the twin
0 / 0
Place—
Click or drag on the table to put the block (or the bin) there. MuJoCo drops it and the arms can push it around.