CustomersOur own use

How to fix every Sentry issue with an agent, without a person on call

A Sentry alert posts to a hosted URL. The machine collects it, the skill finds the cause, lands a small tested fix through a pull request, checks it in production, and asks a person only for what is theirs.

Based on: The self-heal loop of this product, running on the repository behind skillhook.dev since 2026-10-06.

Trigger
A Sentry alert, and a sweep every three hours
Lands alone
A small fix with a test that failed before it
Asks a person for
Migrations, environment, billing, the loop itself
Hosted URLSchedulesInboxAlertsJobs

The situation

Skillhook Cloud is a Next.js application on Vercel with Sentry on it. A production error used to mean that someone noticed it, found the cause, wrote the fix, waited for the deployment and the smoke test, and resolved the issue; or that nobody did, until the next time. The repository already had what an agent needs to do that well: AGENTS.md with hard rules, pnpm check, a smoke test that runs against production after every deployment, and a ruleset on main that wants a pull request with green checks on an up-to-date branch.

What was missing was the part that starts when Sentry notices something. Since 2026-10-06 that part is a hosted URL on this service, a mac mini and the repository's self-heal skill. Until then a claude.ai routine ran a sweep every three hours and a person merged what it proposed; this loop replaces it. This page is the loop as it runs on skillhook.dev, with a one-file version you can install on your own repository.

How it runs

  1. Sentry. An alert rule on the project fires for a new or regressed production error, at most once a day per issue. It has two actions: a GitHub issue labelled sentry and self-heal (the record; the fix's Fixes #N closes it), and the internal integration "Skillhook self-heal", whose webhook URL is the hook's hosted URL on Skillhook Cloud.
  2. The hosted URL. The machine has no public URL. The cloud accepts the delivery, seals it, and hands it over on the machine's next sync (it would wait up to 72 hours). The machine verifies Sentry-Hook-Signature with the integration's client secret, which never left its .env.
  3. skillhook. The hook skillhook-cloud-self-heal is declared in the repository's skillhook.yaml, served because the checkout is linked with skillhook link, and names .claude/skills/self-heal/SKILL.md. Its when filters keep it to the alert rule's action (sentry-hook-resource: event_alert, action: triggered) and to our project; its schedule (17 */3 * * *, UTC) runs the sweep every three hours; one job at a time, two hours at most.
  4. The job. Claude Code (Opus) in a worktree of its own: the root cause, a test that fails before the fix, the fix, pnpm check and the build, a pull request labelled self-heal with auto-merge (squash) when the rules allow it, then the production deployment and the smoke run after it, the Sentry issue resolved with a note pointing at the pull request, and a follow-up pull request with a guard against the class of bug. Its progress reports name the job (the Sentry short id and the error, or "Sweep") and the step: understand, fix, land, verify, prevent.
  5. The inbox. Each run under its title with the step it is on, then the result: a headline, a summary with how to check it, and links to the Sentry issue, the pull requests, the deployment and the smoke run. A run that needs a person comes first, its options as buttons, and the needs_human alert goes to the organisation's channels.

Set it up with your agent

The prompt below walks your agent (Claude Code or Codex with the skillhook plugin) through the same setup for your own repository: skillhook on the machine, the GitHub CLI, the Sentry integration, the hosted URL, the alert rule and a drill event. It never asks you for a secret; you run skillhook secret set in your own terminal. Paste the SKILL.md from the next section as your second message.

The skill

Ours is two files in the repository: the hook (authentication, filters, schedule, timeout) in skillhook.yaml, served with skillhook link, and the skill (runner, model, budget, the body) in .claude/skills/self-heal. The file below is the same loop as one installable skill, with our repository's names replaced by placeholders to edit. Every field in the frontmatter is one skillhook validates, and the job API it names (job_progress, response.json) is what every run has.

Where a person comes in

Nothing deploys to production without a person's yes. Every change goes through a pull request and the ruleset on main, and the loop may turn on auto-merge only for a fix whose root cause it understands and proves with a test that failed before it. The rest is a person's: anything under supabase/ (a migration reaches production before the code that uses it, and only a person applies it); an environment variable or any other change to Vercel, Supabase, Sentry, Stripe or AgentMail; anything that could weaken a hard rule, or touches authentication, API keys, OAuth, cryptography, billing or what the error scrubber lets through; the loop itself, how production is deployed, or a lint rule, test or check that would be loosened; a new runtime dependency; more than about 300 changed lines; a second attempt at an error a self-heal fix already claimed; a cause it could not establish.

For those it opens the pull request without auto-merge, labels it needs-human, writes its evidence and one precise question on GitHub, and ends the run with the outcome needs_human: the question as the headline, the alternatives as options, one recommended. The inbox shows them as buttons; the needs_human alert reaches Slack, a signed webhook or email; an answer in the inbox resumes the agent's session, and an answer on GitHub reaches the next sweep. It never waits inside the run (job_ask_human is not used): one job at a time means a waiting run would hold back every alert behind it.

What to check after

  • Deliveries. The alert's delivery, accepted, via the hosted URL. A rejected one names the reason (a wrong client secret is invalid_signature); the integration's other hooks and other projects' alerts show as skipped, which is the filters doing their job.
  • The job. In Jobs and in the Inbox: its title, its step while it runs, its cost, the headline and the links when it ends; stdout is the whole transcript when something looks wrong.
  • Replay once a cause is fixed: a delivery with replay_delivery (the Sentry secret, the token's permissions, a gh login that expired), a job with replay_job. The skill checks history before it acts, so a replay never opens a second pull request for the same error.
  • GitHub. Pull requests labelled self-heal; issues and pull requests labelled needs-human are the questions waiting for a person.
  • A rehearsal. One synthetic production event with a fingerprint of its own is a new issue: the alert fires, the hosted URL takes it, the run closes it as not a defect. In our rehearsal on 2026-10-06 the event reached the agent 16 seconds after it was sent.

Hosted webhook URLs · Jobs and answers · Alerts