Post #888 177 May 11, 2025, 00:56 UTC turns out pre-training on 4chan is an important part of aligning the LLMshttps://arxiv.org/abs/2505.04741 arXiv.org When Bad Data Leads to Good Models In large language model (LLM) pretraining, data quality is believed to determine model quality. In this paper, we re-examine the notion of "quality" from the perspective of pre- and post-training... 👍 1 😱 1