TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1880 201
How to Evaluate Consistency?

Annotation usually goes through strict stages: data collection → guide creation → annotation launch usually with overlap → label aggregation → model training → quality evaluation.

Let's say we want to get data annotations for a sentiment classifier of support conversations. We have obtained labels from three annotators for some number of real dialogues.

🟢 How can we understand how high-quality the annotation result is? The Agreement

In such a situation, we can calculate the agreement coefficient, The Agreement. It shows the proportion of matching labels between annotators.

In our case, the average pairwise agreement is 83.9%, which is quite good. However, there is a catch: this coefficient, like accuracy in the case of binary classification, can be misleading in the case of class imbalance. In our dataset, more than 70% of the labels fall on two classes — "Neutral" and "Confusion." Let's use other statistical coefficients to ensure high consistency:

📌 Cohen's Kappa (Cohen's Kappa) : pairwise coefficient. It evaluates the normalized agreement between two annotators.

How to interpret the results:

<0.6 : poor consistency.
0.6...0.8 : good consistency, can be used in practical tasks.
>0.8 : very high consistency.

📌 Fleiss's Kappa (Fleiss's Kappa) — suitable for evaluating consistency between several (more than two) annotators.

In our case, Cohen's Kappa varies from 0.7 to 0.84, indicating high consistency. For additional verification, we took several hundred random examples from the dataset and manually assigned labels. It turned out that in 44% of cases, our labels and annotator labels differed that is, annotators on average consistently assign incorrect labels. That is why even high consistency scores do not guarantee high-quality data annotation.


•••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →