Dyna just dropped a world model for robotics achieving a new record for scale of pretraining using human video data - 1 million hours This is 2 orders of magnitude larger than NVIDIA's Egoscale dataset (~20k hours) They established scaling laws parallel to the scaling laws identified in GPT-3. Pretraining on more human video data increased robot performance and generalization across 4 orders of magnitude Most importantly, the scaling fit power laws with no plateau
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours, • this human data scaling law implied a scaling law on never seen robot data, • both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge 🧵
· 58K Views