Rubric
A rubric is a written, testable definition of output quality — the dimensions being judged, the scale for each, and worked examples at each level — used so that two different graders reach the same score.
Writing the rubric is usually the highest-value part of an evaluation engagement, because it is the first time a team is forced to agree on what "good" means. Persistent disagreement between graders is not a people problem; it is a sign the rubric is ambiguous on a specific dimension, and fixing that ambiguity is the deliverable.
Related terms
- Inter-rater agreement
Inter-rater agreement measures how consistently two or more independent graders assign the same score to the same content, and is the standard check on whether a quality rubric is applied reliably.
- LLM-as-judge
LLM-as-judge is an evaluation technique where one language model scores another model’s output against a written rubric, replacing human graders for tasks where output quality cannot be checked by exact match.
Tell us what you are building
Send us the problem in a paragraph. You will get a straight answer on whether we can help, and what we would do first.
Book a call