🟢 How one AI taught another?
The meta-model selects hyperparameters and algorithms used to train the base model During training. It turns out that the training evolves, and the system learns how to learn better 👥
Here, this idea was taken and applied to RL. Technically, there are two levels of trainable parameters. The first is the usual policy of our agent. The second is the meta-parameters that determine the rule by which the policy will be updated.
To optimize the meta-parameters, we run many agents with different policies in different environments. Their experience is the data for training the meta-model. The more data it sees, the better the update rule becomes and, consequently, the more efficiently it trains the agents.
With this approach, the authors managed to synthesize a training algorithm that outperformed previous human solutions. On the Atari game benchmark, an agent trained with its help scored a hundred.
Of course, such achievements require a sea of computing power + it’s not guaranteed that if it works in one area, it will work in another. But interesting, interesting.
And by the way, is this already singularity? 😛
🤖 Data Science, ML & Big Data with @DataXplore
