# Flaw families

> The kinds of flaw an audit finds. Some flaws inflate a score and some deflate it. Only a material flaw makes a task flawed.

Source: https://docs.horizon-audit.online/docs/reference/flaw-families

A task is flawed when its score does not measure the work. A flaw either inflates a score, so that an agent is paid for work it did not do, or deflates a score, so that an agent is not paid for work it did.

## Flaws that inflate a score

| Kind of flaw                              | What it means                                                                          |
| ----------------------------------------- | -------------------------------------------------------------------------------------- |
| The score can be had without the work     | An agent that does nothing, or that fakes its output, gets full reward.                |
| The answer is in the environment          | The agent can read the answer, for example from the history or from the image.         |
| The tests check less than the task asks   | A wrong answer passes, because the grader does not test what the instruction requires. |
| The task does not test the skill it names | An agent can pass without the skill that the task claims to measure.                   |

## Flaws that deflate a score

| Kind of flaw                                | What it means                                                                         |
| ------------------------------------------- | ------------------------------------------------------------------------------------- |
| The tests and the instruction disagree      | A right answer fails, because the grader asks for something the instruction does not. |
| The task cannot be run or reviewed          | A part is missing, the environment does not build, or a requirement cannot be met.    |
| Grading depends on something that can break | The result rests on something outside the task that can change or stop.               |

A task that misses a part is not counted as flawed, and it is not counted as clean. It comes back as "Needs from you".

## Proved flaws and flags

A proved flaw is shown in a run. It comes with the saved output of that run.

A flag is something worth a look. A flag never decides a verdict.

## What material means

A flaw is material when it can change a realistic agent's score. Only a material flaw makes a task flawed. A flaw is material in these cases:

- A realistic answer is scored the wrong way, whether the answer is right or wrong.
- The instruction contradicts the grader.
- The answer key is wrong.
- The reference solution fails.
- The answer leaks to the agent.
- A requirement cannot be met.

A defect that exists but cannot change a realistic score is not material. It stays on the record as a proved flaw with its finding. The task can still be clean. The result then shows the flaw as a flag, and it gives the reason why the flaw is not material.

## Who stands behind a verdict

The verdict is computed from the saved output of runs. An engineer confirms every flawed verdict before you see it. You never see a draft.

## What you get for a flaw

Each flaw of a flawed task has a finding. A finding has three parts:

- where the flaw is, with the lines quoted
- how it was established, as numbered steps you can repeat
- what the finding does not claim

A finding does not say how to change the task. See [read a verdict](/docs/guide/read-a-verdict).
