😎🔥A small collection of useful datasets:
Synthia-v1.5-I – a dataset that includes over 20,000 technical questions and answers. It uses system prompts in the Orca style to generate diverse responses, making it a valuable resource for training and testing LLMs on complex technical data.
HelpSteer2 – an English-language dataset designed for training reward models that improve the utility, accuracy, and coherence of responses generated by other LLMs.
LAION-DISCO-12M – includes 12 million links to publicly available YouTube tracks with metadata. The dataset is created to support research in machine learning, sound processing model development, musical data analysis, audio data processing, and training recommender systems and applications.
Universe – a large-scale collection containing astronomical data of various types: images, spectra, and light curves. It is intended for research in astronomy and astrophysics.
Post #768
864

- ❤ 1