Skip to content
Evaluation

Inter-rater agreement

Inter-rater agreement measures how consistently two or more independent graders assign the same score to the same content, and is the standard check on whether a quality rubric is applied reliably.

It is reported with statistics such as Cohen’s kappa or Krippendorff’s alpha, which correct for agreement that would occur by chance. Low agreement invalidates every downstream number: if humans cannot apply the rubric consistently, a model grader trained or validated against it inherits the same noise.

Related terms

  • Rubric

    A rubric is a written, testable definition of output quality — the dimensions being judged, the scale for each, and worked examples at each level — used so that two different graders reach the same score.

  • LLM-as-judge

    LLM-as-judge is an evaluation technique where one language model scores another model’s output against a written rubric, replacing human graders for tasks where output quality cannot be checked by exact match.

← All terms

Tell us what you are building

Send us the problem in a paragraph. You will get a straight answer on whether we can help, and what we would do first.

Book a call