What Horizon Audit is
Horizon Audit checks benchmark tasks before release. Each task comes back clean, flawed, or with a request for what is missing.
Horizon Audit checks benchmark tasks before they are released. You submit Harbor tasks, and each comes back with a verdict: clean, for the exact files that were audited, or flawed, with the flaw and its proof.
A task is flawed when its score does not measure the work. The common cases are these:
- An agent that does nothing passes.
- The reference solution fails.
- The tests check less than the instruction asks.
- The answer can be read from the environment.
- The grading depends on something that can break.
Who it is for
Horizon Audit is for the people who prepare tasks and release them:
- Dataset vendors who prepare tasks for model labs.
- Benchmark publishers.
- Teams that turn workflows into environments for agents.
Model labs use the verdicts to decide which tasks to trust.
What you send
You send Harbor tasks. Harbor is the input format. A Harbor task has five parts:
| Part | File |
|---|---|
| An instruction | instruction.md |
| A configuration | task.toml |
| An environment | environment/Dockerfile, or one of two other forms |
| A reference solution | solution/solve.sh |
| A verifier | tests/test.sh |
You send the tasks as one archive of task folders. Prepare a Harbor task says what each part must hold.
What you get back
Each task comes back with one of three labels.
| Label | What it means |
|---|---|
| Clean | Every promise of a clean verdict held. The verdict holds for the exact files that were audited. |
| Flawed | A flaw that can change a realistic score was proved. You get the flaw, where it is, and the proof. |
| Needs from you | The task cannot be audited until you supply something. This is not a verdict, and it is never counted as clean. |
A result appears as soon as its task is done. Tasks do not wait for each other.
What a clean verdict proves
- No reward without the work. An agent that does nothing, or only looks busy, earns nothing.
- The task can be solved as written. The task's own solution earns full reward in the task's own environment.
- A wrong answer is refused. An answer that is close but wrong does not pass.
- A right answer is accepted. A correct answer passes, also when it does not look like the author's.
- The answer cannot be found without the work. The answer is not there to be read from the environment.
Anything else worth a look is listed beside the verdict as a flag. A flag never decides a verdict.
The principles
- Not checked is never clean.
- Clean is a claim about what was tested.
- Tasks are never edited to make them run.
- You never see a draft.
The verdict is computed from the saved output of runs. An engineer confirms every flawed verdict before you see it. You never see a draft.
Access
Access is by invitation, and our team creates the account. Write to contact@horizonanalyticslabs.com.
Where to go next
- After you submit: what happens to a task, and what you see.
- Prepare a Harbor task: a checklist to follow before you submit.
- Submit tasks: the archive, the preview, and the confirmation.
- The intake standard: the nine rules a task must meet before it can be audited.
- Read a verdict: what a clean result and a flawed result contain.
- Fix and resubmit: how a request is met, and how a changed task is sent again.