📋💡🔎Selection of datasets for NLP
КартаСловСент — words and expressions equipped with a tonal label (“positive”, “negative”, “neutral”) and a scalar value of the strength of the emotional-evaluative charge from the continuous range [-1, 1].
WikiQA is a set of pairs of questions and proposals. They were collected and annotated to explore answers to questions in open domains
Amazon Reviews dataset - this dataset consists of several million Amazon customer reviews and their ratings. The dataset is used to enable fastText to learn by analyzing customer sentiment. The idea is that despite the huge volume of data, this is a real business problem. The model is trained in minutes. This is what sets Amazon Reviews apart from its peers.
Yelp dataset is a set of businesses, reviews and user data that can be used in a Pet project and scientific work. Yelp can also be used to train students while working with databases, learning NLP, and as a sample of manufacturing data. The dataset is available as JSON files and is a “classic” in natural language processing.
Post #606
2.21K