Policy-Bound Evaluation
Evaluation scored against explicit plans, tasks, evidence, and policies.
Policy-Bound Evaluation is the more specific form of Structured Judgment, emphasizing that evaluation criteria should come from existing policies and standards — not ad hoc assessments. The team's coding standards, security policies, documentation requirements, and testing protocols become the evaluation rubric. This ensures consistency: the same code is evaluated the same way regardless of who (or what) reviews it. It also ensures alignment: the evaluation measures what the organization actually cares about.
More in Evaluation
Structured Judgment
Evaluation where an AI judge scores work against explicit plans, tasks, evidence, and policies rather than subjective impressions.
Eval Drift
Gradual misalignment between what an evaluation measures and what actually matters.
Judgment Bias
Systematic skew in AI evaluation due to unexamined assumptions, prompt framing, or training artifacts.
Rubric Rot
Decay in evaluation criteria relevance over time — the eval no longer tests what it should.
Eval Capture
When an agent optimizes for passing evaluation rather than doing the actual work.
Eval Integrity
Evaluation that remains aligned with real-world outcomes over time — resistant to Eval Drift, Rubric Rot, and Eval Capture.