Why does it make sense to spend Less Time On Modeling & Experiments and More Time On Data Preparation in situations With Low Metrics?
Suppose we have an emotion classifier and want to boost its metrics.
☞ Change the training approach : experiment with architecture, pretraining, optimizers.
☞ Work with the data : check the dataset, review the labeling, find noise and errors.
☞ Or completely abandon the classic ML model and try feeding everything to an LLM, hoping for the model's zero/few-shot capabilities.
Most engineers choose the first two options. But, as practice shows, it is precisely the search for errors in datasets and improving data quality that gives the most noticeable gain.
Surface defect detection task Example
If you improve model or training approach, you'll not notice quality improvement, or it'll be minimal.
If you work with data, you will see a significant quality increase.
🤖 Data Science, ML & Big Data with @DataXplore
