Flaw families
The kinds of flaw an audit finds. Some flaws inflate a score and some deflate it. Only a material flaw makes a task flawed.
A task is flawed when its score does not measure the work. A flaw either inflates a score, so that an agent is paid for work it did not do, or deflates a score, so that an agent is not paid for work it did.
Flaws that inflate a score
| Kind of flaw | What it means |
|---|---|
| The score can be had without the work | An agent that does nothing, or that fakes its output, gets full reward. |
| The answer is in the environment | The agent can read the answer, for example from the history or from the image. |
| The tests check less than the task asks | A wrong answer passes, because the grader does not test what the instruction requires. |
| The task does not test the skill it names | An agent can pass without the skill that the task claims to measure. |
Flaws that deflate a score
| Kind of flaw | What it means |
|---|---|
| The tests and the instruction disagree | A right answer fails, because the grader asks for something the instruction does not. |
| The task cannot be run or reviewed | A part is missing, the environment does not build, or a requirement cannot be met. |
| Grading depends on something that can break | The result rests on something outside the task that can change or stop. |
A task that misses a part is not counted as flawed, and it is not counted as clean. It comes back as "Needs from you".
Proved flaws and flags
A proved flaw is shown in a run. It comes with the saved output of that run.
A flag is something worth a look. A flag never decides a verdict.
What material means
A flaw is material when it can change a realistic agent's score. Only a material flaw makes a task flawed. A flaw is material in these cases:
- A realistic answer is scored the wrong way, whether the answer is right or wrong.
- The instruction contradicts the grader.
- The answer key is wrong.
- The reference solution fails.
- The answer leaks to the agent.
- A requirement cannot be met.
A defect that exists but cannot change a realistic score is not material. It stays on the record as a proved flaw with its finding. The task can still be clean. The result then shows the flaw as a flag, and it gives the reason why the flaw is not material.
Who stands behind a verdict
The verdict is computed from the saved output of runs. An engineer confirms every flawed verdict before you see it. You never see a draft.
What you get for a flaw
Each flaw of a flawed task has a finding. A finding has three parts:
- where the flaw is, with the lines quoted
- how it was established, as numbered steps you can repeat
- what the finding does not claim
A finding does not say how to change the task. See read a verdict.