YES, but wisely.
At first look, it seems everything is simple: we write a prompt, send it to the model, and get the result. It's cheaper and faster than human annotation.
But in practice, to make the prompt work effectively, it requires many iterations, improvements, and experiments. For example, for our emotion model, we spent 3 weeks optimizing the prompt. As a result, we obtained a good correlation coefficient between the LLM and assessors — on average, 0.81.
Then we trained two classifiers on different datasets:
📌 Exclusively with LLM labels.
📌 With assessor labels after the work done to improve consistency and reduce errors.
LLM annotation — Weighted F1 0.66
Assessor annotation — Weighted F1 0.7
As a result, the quality of the LLM labels is only slightly below the reference annotation. At the same time, we saved a large amount of time and human resources.
Currently, LLMs can be considered as junior assessors. The model responds approximately like a human, and in the absence of resources, it can be used for data annotation. In some tasks, this will yield quality comparable to the reference annotation.
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
