?
- What is happening in specialized topics like audio/video/photo analysis and generation? Can the models write down the score of a symphony looking at the video (with sound), or find an error in the performance of a Beethoven's sonata? How do they do it (what is replacing the tokens etc.)?
If all goes well, we can have some live experiments and comments (a session of magic and its exposure).
➰ ВК
Post #317
266