Rubric Rot
Decay in evaluation criteria relevance over time — the eval no longer tests what it should.
Rubric Rot is a specific mechanism of Eval Drift focused on the rubric itself. Individual criteria that were once meaningful become meaningless: a check for "uses modern syntax" meant something when the team was migrating from legacy patterns, but a year later it's always true and provides no signal. Criteria accumulate without cleanup: new checks are added but old irrelevant ones are never removed. The rubric grows into a checklist that everyone passes but that catches nothing, creating a false sense of quality assurance.
More in Evaluation
Structured Judgment
Evaluation where an AI judge scores work against explicit plans, tasks, evidence, and policies rather than subjective impressions.
Policy-Bound Evaluation
Evaluation scored against explicit plans, tasks, evidence, and policies.
Eval Drift
Gradual misalignment between what an evaluation measures and what actually matters.
Judgment Bias
Systematic skew in AI evaluation due to unexamined assumptions, prompt framing, or training artifacts.
Eval Capture
When an agent optimizes for passing evaluation rather than doing the actual work.
Eval Integrity
Evaluation that remains aligned with real-world outcomes over time — resistant to Eval Drift, Rubric Rot, and Eval Capture.