Fixed how results are scored and published
Published results now come from a single CI run, measured under one base prompt on one day, and the score counts scenarios completed rather than checks passed.
Compare agent performance across Hookdeck features: queueing, alerts, retries, filters, and more. Agents work in real Hookdeck projects using the CLI and a live API key.
| Agent/Avg. Performance | Build | Troubleshoot | Recover |
|---|---|---|---|
| 92% | 100% | 100% | |
| 85% | 100% | 100% | |
| 69% | 100% | 100% |
Results from , run 33484972784, published in v0.4.0.
Published results now come from a single CI run, measured under one base prompt on one day, and the score counts scenarios completed rather than checks passed.
Outpost had one scenario in this benchmark. It now has five, covering customer subscriptions, a destination switched off after repeated failures, alerting, delivery to a queue, and narrowing what one customer receives.
An agent working in a sandbox has no terminal, so hookdeck listen cannot prompt it to log in. hookdeck ci reads the key already in the environment and does — but the skill never said so, and a weak model that hit the interactive path concluded its API key had been rejected, then documented a local setup instead of building one.