# What Horizon Audit is

> Horizon Audit checks benchmark tasks before release. Each task comes back clean, flawed, or with a request for what is missing.

Source: https://docs.horizon-audit.online/docs

Horizon Audit checks benchmark tasks before they are released. You submit Harbor tasks, and each comes back with a verdict: clean, for the exact files that were audited, or flawed, with the flaw and its proof.

A task is flawed when its score does not measure the work. The common cases are these:

- An agent that does nothing passes.
- The reference solution fails.
- The tests check less than the instruction asks.
- The answer can be read from the environment.
- The grading depends on something that can break.

## Who it is for

Horizon Audit is for the people who prepare tasks and release them:

- Dataset vendors who prepare tasks for model labs.
- Benchmark publishers.
- Teams that turn workflows into environments for agents.

Model labs use the verdicts to decide which tasks to trust.

## What you send

You send Harbor tasks. Harbor is the input format. A Harbor task has five parts:

| Part                 | File                                                |
| -------------------- | --------------------------------------------------- |
| An instruction       | `instruction.md`                                    |
| A configuration      | `task.toml`                                         |
| An environment       | `environment/Dockerfile`, or one of two other forms |
| A reference solution | `solution/solve.sh`                                 |
| A verifier           | `tests/test.sh`                                     |

You send the tasks as one archive of task folders. [Prepare a Harbor task](/docs/guide/prepare-a-task) says what each part must hold.

## What you get back

Each task comes back with one of three labels.

| Label          | What it means                                                                                                   |
| -------------- | --------------------------------------------------------------------------------------------------------------- |
| Clean          | Every promise of a clean verdict held. The verdict holds for the exact files that were audited.                 |
| Flawed         | A flaw that can change a realistic score was proved. You get the flaw, where it is, and the proof.              |
| Needs from you | The task cannot be audited until you supply something. This is not a verdict, and it is never counted as clean. |

A result appears as soon as its task is done. Tasks do not wait for each other.

## What a clean verdict proves

- **No reward without the work.** An agent that does nothing, or only looks busy, earns nothing.
- **The task can be solved as written.** The task's own solution earns full reward in the task's own environment.
- **A wrong answer is refused.** An answer that is close but wrong does not pass.
- **A right answer is accepted.** A correct answer passes, also when it does not look like the author's.
- **The answer cannot be found without the work.** The answer is not there to be read from the environment.

Anything else worth a look is listed beside the verdict as a flag. A flag never decides a verdict.

## The principles

- Not checked is never clean.
- Clean is a claim about what was tested.
- Tasks are never edited to make them run.
- You never see a draft.

The verdict is computed from the saved output of runs. An engineer confirms every flawed verdict before you see it. You never see a draft.

## Access

Access is by invitation, and our team creates the account. Write to contact@horizonanalyticslabs.com.

## Where to go next

- [After you submit](/docs/how-an-audit-runs): what happens to a task, and what you see.
- [Prepare a Harbor task](/docs/guide/prepare-a-task): a checklist to follow before you submit.
- [Submit tasks](/docs/guide/submit): the archive, the preview, and the confirmation.
- [The intake standard](/docs/guide/intake-standard): the nine rules a task must meet before it can be audited.
- [Read a verdict](/docs/guide/read-a-verdict): what a clean result and a flawed result contain.
- [Fix and resubmit](/docs/guide/fix-and-resubmit): how a request is met, and how a changed task is sent again.
