Induction introduced imagination models: a new foundation model architecture that unlocks learning from internet-scale video.
The first imagination model, Photon-1, learned to use a computer by watching 18 years of screen recording video without action labels.
Imagination models predict what happens next in a video in representation space.
Predicting future states teaches Photon-1 to use a computer, implicitly. After a small finetune to teach it to use the keyboard and mouse, Photon-1 can imagine the future and act to get there.
Despite only ever seeing computer video, Photon-1 learns general world concepts.
With a small amount of finetuning, it can predict billiard physics and checkers states better than baselines trained on the same data.
Photon-1 also picked up human habits from its pretraining video.
It learned to use AI tools the way people do. Sometimes it prompts ChatGPT, checks the output, and steers the LLM until the task is done.
Post #4379
640