AI & LLMs

Inter-Annotator Agreement

also: inter-expert agreement

How often two humans doing the same labelling task give the same answer — the practical ceiling on any model’s measured accuracy.

Inter-annotator agreement is measured by having two or more domain experts label the same cases independently and counting how often they match. It is the cheapest and most under-used measurement in enterprise AI, because it converts an argument about accuracy into a number everyone can see.

Worked example: if two senior claims adjusters agree on 72% of cases, no model is going to be judged “95% accurate” on the same task, and the disagreements they had are precisely the rules nobody had written down. Gotcha: measure it before an accuracy target goes into a contract; afterwards it looks like an excuse rather than a finding.