TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1834 184
Teaching models to see like humans

Several vision models tested on the classic game "Find the odd picture out"? and found that they don't always reason like humans because models, still see the world differently than humans.

This is an important drawback because if a model does't structure its internal representation of world like a human, its decisions can be illogical and unreliable.

🟢 What DeepMind did?
Humans group objects by meaning, while models more often group by visual/textural features. For example, in the task <starfish, cat, fox>, humans choose the starfish because it lives in water, while models choose the cat because the picture stands out in the color palette.

– Took a large visual model, froze it, and artificially attached a small adapter trained only on a dataset with human answers for odd-one-out.

– Generated a huge dataset with millions of odd-one-out decisions using this model.

– Based on this large dataset, fine-tuned other models so that their internal representations became closer to human logic of grouping.

So, it seems like the models were just trained on some children's game, nothing surprising. But it turned out that this qualitatively changed their hidden space. See the third screenshot: on the left, the model before alignment, on the right, after. You can see clear clusters appearing that correspond to human logic (for example, animals, food, etc.). Beautiful, right?

Also, such a model turned out to be more robust in terms of distribution shifts. For example, if the background, lighting, or other conditions in the pictures change, its answers still follow logic.


Overall, it can be expected that such a model will generalize faster than usual.

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →