The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions.
He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a ...