⚡️🔎Fully Synthetic Dataset
A huge dataset consisting entirely of synthetic data has appeared on Hugging Face.
The LLM (in this case GPT-4o + VLLM) generates answers by representing itself each time with some character: for example, a chemical scientist or a musician.
Synthetic data can sometimes help a lot (especially when the task is abstract and there is no structured information), but they are still treated with caution. They are not realistic enough, they are not diverse enough, and they potentially harbor hallucinations. It is still unclear whether we will ever be free to use “synthetics”, but it is actively being worked on.
Post #703
1.54K