Horizon Audit
Reference

Flaw families

The kinds of flaw an audit finds. Some flaws inflate a score and some deflate it. Only a material flaw makes a task flawed.

A task is flawed when its score does not measure the work. A flaw either inflates a score, so that an agent is paid for work it did not do, or deflates a score, so that an agent is not paid for work it did.

Flaws that inflate a score

Kind of flawWhat it means
The score can be had without the workAn agent that does nothing, or that fakes its output, gets full reward.
The answer is in the environmentThe agent can read the answer, for example from the history or from the image.
The tests check less than the task asksA wrong answer passes, because the grader does not test what the instruction requires.
The task does not test the skill it namesAn agent can pass without the skill that the task claims to measure.

Flaws that deflate a score

Kind of flawWhat it means
The tests and the instruction disagreeA right answer fails, because the grader asks for something the instruction does not.
The task cannot be run or reviewedA part is missing, the environment does not build, or a requirement cannot be met.
Grading depends on something that can breakThe result rests on something outside the task that can change or stop.

A task that misses a part is not counted as flawed, and it is not counted as clean. It comes back as "Needs from you".

Proved flaws and flags

A proved flaw is shown in a run. It comes with the saved output of that run.

A flag is something worth a look. A flag never decides a verdict.

What material means

A flaw is material when it can change a realistic agent's score. Only a material flaw makes a task flawed. A flaw is material in these cases:

  • A realistic answer is scored the wrong way, whether the answer is right or wrong.
  • The instruction contradicts the grader.
  • The answer key is wrong.
  • The reference solution fails.
  • The answer leaks to the agent.
  • A requirement cannot be met.

A defect that exists but cannot change a realistic score is not material. It stays on the record as a proved flaw with its finding. The task can still be clean. The result then shows the flaw as a flag, and it gives the reason why the flaw is not material.

Who stands behind a verdict

The verdict is computed from the saved output of runs. An engineer confirms every flawed verdict before you see it. You never see a draft.

What you get for a flaw

Each flaw of a flawed task has a finding. A finding has three parts:

  • where the flaw is, with the lines quoted
  • how it was established, as numbered steps you can repeat
  • what the finding does not claim

A finding does not say how to change the task. See read a verdict.

On this page