TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1887 443
What to do if the annotator markup does not match the expert markup?

Consistency co-efficients do not always reflect the final quality of the markup and the model. almost half of the examples are marked incorrectly, the annotators agree with each other but not with the experts.

🟢 How can we improve the markup?

CHECK the task formulation and specify a detailed guide with corner cases. You can take a sample, markup it according to the guide, and see where disputes arise, these places need to be clarified.

COLLECT a test dataset with gold markup using an expert. After that, you can select annotators with high performance on the test set or hold a briefing meeting with all annotators to discuss the errors.

SPLIT the work into chunks and add a golden set for validation in each. This will allow you to evaluate the quality of the markup iteratively and monitor how much the annotators match the golden set.

After implementing these steps, our emotion model's weighted F1 score increased from 0.61 to 0.7, and the discrepancy between expert and annotator markup decreased from 44% to 18%. Also, some small problematic classes improved well:

Gratitude — 0.8 → 0.76
Neutral — 0.7 → 0.75
Satisfactory — 0.68 → 0.74
Impatience — 0.57 → 0.53
Disappointment — 0.57 → 0.55
Confusion — 0.46 → 0.7

Important: low consistency does not always mean poor performance by the annotators. The reasons may be:

Unclear task: it may imply some uncertainty. For example, this often occurs when preparing dialog data for LLMs.

Different backgrounds of the annotators: internal AI trainers and external contractors may understand the task differently. This leads to significant differences in evaluations.


Therefore, ML engineers and data scientists must carefully read the data and understand how it is marked.

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →