If you start fine-tuning an LLM on low-quality social media data (short, popular, clickable posts), it begins to lose its cognitive abilities. Much like how a person loses attention and memory when they doomscroll too much.
🟢 but Why this happens from a technical perspective?
They took Llama 3 8B Instruct and began fine-tuning it on (a) short and very popular posts with many likes, retweets, and replies; and (b) content with low semantic value: clickbait, conspiracy theories, and the like. After that, they measured metrics and compared them with the results before fine-tuning. The results:
– Reasoning quality dropped from 74.9 to 57.2
– Understanding of long context dropped from 84.4 to 52.3
– Alignment tests revealed that the model developed narcissism, Machiavellianism, and psychopathy
Even after additional tuning on clean data, the degradation did not completely disappear.
But the thing is, there is no groundbreaking discovery here. It is all explained by a simple distribution shift. When fine-tuning on short, popular, emotionally charged tweets, the model sees a completely different statistical landscape than during the original pretraining on books, articles, etc.
This shifts the distribution in the embedding space and changes attention patterns. The model constantly sees short texts without logical chains, and naturally, attention masks start to focus more on the last few tokens and lose long-term dependencies that previously ensured high-quality CoT.
Gradient dynamics also work against us here. The loss is simply minimized by superficial correlations, and parameters responsible for long causal relationships barely get updated. So the model loses the ability to reason at length. The authors call this phenomenon thought-skipping.
That's it. Just another proof that data is everything. Now you can go back to scrolling reels ☕️
🤖 Data Science, ML & Big Data with @DataXplore
