# How well do AI agents build with Hookdeck?

Compare agent performance across Hookdeck features: queueing, alerts, retries, filters, and more. Agents work in real Hookdeck projects using the CLI and a live API key.

Results from 1 September 2026. The workflow run behind these numbers is [33484972784](https://github.com/hookdeck/evals/actions/runs/33484972784).

348 runs recorded · 14 findings filed · 2 fixes shipped

Cells are `passed/run`. A dash means that pairing has not been run. Fractions rather than percentages throughout: `2/2` is a weaker claim than `100%` implies, and the difference matters when a stage holds two scenarios and another holds eleven.

## By eval

| Eval | Claude Code Sonnet 5 (no skills) | Claude Code Sonnet 5 (+skills) | Codex GPT-5.6 (no skills) | Codex GPT-5.6 (+skills) | Codex GPT-5.4-mini (no skills) | Codex GPT-5.4-mini (+skills) | Total |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Delivery alerts | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Duplicate events | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Slow consumer | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Rate limited endpoint | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Enterprise orders | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| High value retries | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 0/1 | 4/6 |
| Failing deliveries | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Partial outage | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Listen locally | 0/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 5/6 |
| Customer subscriptions | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 1/1 | 5/6 |
| Disabled destination | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 5/6 |
| Operator events | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Queue destination | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 1/1 | 5/6 |
| Topic scoping | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Paused connection | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Scoped redelivery | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 6/6 |
| Reshape payload | 1/1 | 1/1 | 1/1 | 1/1 | 1/1 | 0/1 | 5/6 |
| Stripe express | 0/1 | 1/1 | 1/1 | 0/1 | 0/1 | 0/1 | 2/6 |
| Elevenlabs callbacks | 1/1 | 1/1 | 0/1 | 1/1 | 1/1 | 0/1 | 4/6 |

## By agent

| Agent | Build | Troubleshoot | Recover | Total |
| --- | --- | --- | --- | --- |
| Claude Code Sonnet 5 (no skills) | 11/13 | 2/2 | 4/4 | 17/19 |
| Claude Code Sonnet 5 (+skills) | 13/13 | 2/2 | 4/4 | 19/19 |
| Codex GPT-5.6 (no skills) | 12/13 | 2/2 | 4/4 | 18/19 |
| Codex GPT-5.6 (+skills) | 12/13 | 2/2 | 4/4 | 18/19 |
| Codex GPT-5.4-mini (no skills) | 9/13 | 2/2 | 4/4 | 15/19 |
| Codex GPT-5.4-mini (+skills) | 9/13 | 2/2 | 3/4 | 14/19 |

## By journey

| Journey | Claude Code Sonnet 5 (no skills) | Claude Code Sonnet 5 (+skills) | Codex GPT-5.6 (no skills) | Codex GPT-5.6 (+skills) | Codex GPT-5.4-mini (no skills) | Codex GPT-5.4-mini (+skills) | Total |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Build | 11/13 | 13/13 | 12/13 | 12/13 | 9/13 | 9/13 | 66/78 |
| Troubleshoot | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 12/12 |
| Recover | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 3/4 | 23/24 |

## By product

| Product | Claude Code Sonnet 5 (no skills) | Claude Code Sonnet 5 (+skills) | Codex GPT-5.6 (no skills) | Codex GPT-5.6 (+skills) | Codex GPT-5.4-mini (no skills) | Codex GPT-5.4-mini (+skills) | Total |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Event Gateway | 12/14 | 14/14 | 13/14 | 13/14 | 12/14 | 10/14 | 74/84 |
| Outpost | 5/5 | 5/5 | 5/5 | 5/5 | 3/5 | 4/5 | 27/30 |

## What the experiments are

Each agent runs twice: once with Hookdeck's skills installed and once without, identical in every other respect, so any difference between the two is attributable to the skills. Both arms get the CLI and a live API key, so neither is documentation-only. One model is included deliberately as a weaker baseline, to find the floor beneath the frontier agents rather than to represent typical usage; a low score there is a measurement rather than a product finding.

## Changelog

Not every failure is understood yet. The ones we can explain become findings, and findings that turn out to be product problems become changes. Benchmark entries — repairs to the harness itself — are in the release notes but not listed here.

### v0.4.0 — Fixed how results are scored and published

**Discovered**

- Four of six agents wired a placeholder secret into a live source and reported success — #75

### v0.3.0 — Expanded Outpost coverage to five scenarios

**Discovered**

- Turning on alerts for a disabled destination needs an API that appears in no OpenAPI definition or doc — #34
- Every environment variable on the operator events documentation page is wrong — #32
- A key for the wrong project type is reported as "Not Found" rather than as the wrong key — #39
- The Outpost skill never says you need a key belonging to an Outpost project — #40

### v0.2.0 — Documented no-terminal CLI auth

**Shipped**

- The skill did not say how to authenticate the CLI without a terminal — skills · #27

**Discovered**

- No signal for when a configuration change is in force — #25
- Skills make the weak model worse — #2
- Documentation lookups collapse when skills are installed — #10

## Notes

- Only the benchmark suite is published. A separate regression suite guards incidents already fixed and does not belong on a scoreboard.
- Failures are published deliberately. A benchmark showing only passes would not be worth reading.
- Experiments refresh on different cadences, so two columns can differ in age.
- The table shows one run: the snapshot the latest release points at, not the newest run. Runs happen more often than releases, so newer results may exist that are not shown. Publishing is a deliberate step rather than whatever a scheduled job last produced.
- Every snapshot is kept, so the figures behind a release can be checked afterwards: https://github.com/hookdeck/evals/tree/main/results/runs
- Source, scenarios and scorers: https://github.com/hookdeck/evals
- Built on https://github.com/supabase/evals.
