Documentation

Jobs and answers

Status, outcome and failure kinds; questions agents ask a person; replaying and running skills.

Status and outcome

Every job carries two words. The status says how the process ended; the outcome says whether the task was done, as the agent reported it by writing response.json in its job directory ({outcome, summary, links, data}), or as a structured answer when the skill asks for one. A succeeded run can still need a person; a run that produced nothing is unknown, not failed.

StatusTypeDescription
queued
status
Waiting for a free slot on the machine (two jobs at once by default, one per skill; the queue survives restarts).
running
status
The runner is working; watch_job streams its output.
succeeded
status
The process ended normally. Look at the outcome for whether the task was done.
failed
status
The process ended with an error; the failure kind says which.
timed_out
status
The skill's timeout_seconds passed. The clock is paused while the agent waits for a person.
cancelled
status
Stopped by a person: cancel_job, the dashboard, or skillhook jobs cancel on the machine.
interrupted
status
The server stopped before the run ended.
OutcomeTypeDescription
completed
outcome
The agent reports the task done. A shell command that exits 0 is completed too.
partial
outcome
Some of it was done; the summary says what is missing.
needs_human
outcome
The agent needs a person: it asked a question nobody answered yet, or ended saying so. Listed in the inbox under Needs input.
nothing_to_do
outcome
The event needed no action; a reasoned no-op, not a failure.
failed
outcome
The agent could not do it, or the run did not succeed (every non-succeeded status is failed).
unknown
outcome
The run reported no outcome (no response.json, no structured answer).

The inbox: progress, results and links

The dashboard's Inbox is every job as a person reads it: what it is about, what it needs, how far it got, what came of it and where to look. Its tabs are Needs input (agents waiting for an answer, with their choices as buttons; the default whenever anyone waits), In progress, Done, Failed and All; it sorts by latest activity, by who has waited longest, by progress or by creation, and filters by machine and skill. The webhook delivery that started a job is linked from it.

Agents fill it through the job API skillhook gives every run (skillhook 0.8 or newer; the guardrails ask for all of it):

  • A title, what the job is about in a few words (at most 200 characters), shown instead of the skill's name: job_progress {message, title} or skillhook job progress "…" --title "…", given once.
  • Progress while it runs: a message, a step and a percentage (job_progress {message, step, percent}); the inbox shows the latest with the updates before it and can sort by how far along each job is.
  • The result in one line, the headline (at most 280 characters), with the summary below it: job_set_outcome, skillhook job outcome --headline or response.json.
  • Links with a kind, so they are grouped and labelled: {url, title, kind}, or a bare URL. On the command line, --link "pull_request:[PR #7](https://github.com/acme/api/pull/7)".
  • Choices for a person: options, the recommended one and whether several may be picked (multiple), on a question or on an outcome needs_human. Each becomes a button; one click answers, and a finished run resumes its session with the pick. Several picks reach the agent one per line.

Questions, context, summaries, headlines and progress messages may use Markdown: paragraphs, lists and task lists, emphasis, code and code blocks, links (http, https and mailto only), quotes, tables. It is rendered as text and links, never as HTML, and images are shown as links.

Link kindTypeDescription
source
kind
What started the run: the Sentry issue, the GitHub issue or pull request, the Granola meeting, the Slack thread, the form. Shown first.
pull_request
kind
A pull or merge request the run opened, updated, reviewed or merged.
commit
kind
A commit.
issue
kind
An issue, ticket or task it opened or updated (GitHub, Linear, Jira, Asana).
message
kind
A message it sent or answered (Slack, email, a comment).
document
kind
A document it wrote or changed (Notion, Google Docs, a wiki page, a file).
deploy
kind
A deployment or a preview.
test
kind
How to check the result: a CI run, a preview to try, a test report.
log
kind
Logs, traces or a dashboard it looked at.
result
kind
The result itself when it lives elsewhere: a report, a generated file, a dataset.
other
kind
Anything else. A link without a kind gets one from where it points (a GitHub pull request, a Slack message, a Notion page…).

Agents read the same through list_inbox {view?, sort?, machine?, skill?} (GET /api/v1/inbox, skillhook cloud list_inbox --view needs_input): each job's title, state, progress with its recent updates, headline, summary, links and, while it waits, its question with the choices. get_job and list_jobs carry the title, the headline and the links too.

Failure kinds and what to do

A failed or timed-out job carries failure: {kind, code, retryable, message}, classified on the machine from what the CLI printed. The kind tells you whether the trouble is the account (auth, usage_limit, budget), the moment (rate_limit, crash) or the run itself (max_turns, timeout). list_jobs {failure: ...} filters by it, and get_job shows the message and the progress timeline.

KindWhat it meansWhat to do
authThe runner is not installed or not signed in, or its API key is invalid. When readiness is checked before the run, the job fails at once without starting a process.Run claude login or codex login as the user the service runs as, or put ANTHROPIC_API_KEY / OPENAI_API_KEY in ~/.skillhook/.env. The machine's page and skillhook runners show readiness; a fallback runner named in the skill takes over on its own.
usage_limitThe account's usage limit or quota is reached ("hit your limit", "quota", "limit will reset").Wait for the reset, or name a fallback runner on another account with fallback.on: [usage_limit]; that re-runs a failed job and is safe only for idempotent skills.
rate_limitThe provider answered 429 or "overloaded". Retryable.Replay once it clears. A retry: {attempts, on, backoff_seconds} block in the skill repeats a run on the same runner by itself (rate_limit and crash by default, 30 s apart).
budgetClaude's max_budget_usd was reached before the agent finished.Raise claude.max_budget_usd or narrow the skill; the transcript so far is in stdout.
max_turnsClaude's turn limit was reached.Narrow the task or split the skill; read stdout to see where the agent circled.
not_foundThe runner's command could not start: not on PATH, or the file is gone.Install the CLI where the service can see it; skillhook doctor shows the versions it finds.
timeoutThe skill's timeout_seconds passed (the job's status is timed_out).Raise timeout_seconds in the skillhook: block or make the skill do less. Waiting for a person does not count against it.
crashThe CLI exited with an error or a signal without reporting a result. Retryable.Read stderr (get_job_artifact stderr); the retry block handles the transient ones.
unknownThe run failed and nothing in what the CLI printed matched a known kind.Read stdout and stderr. If skillhook should have recognised it, report_issue tells the Skillhook team.

Questions for a person

An agent is not cut off while it runs. Through a per-run job API (MCP tools injected into the run, or skillhook job ask for shell skills) it can report progress and ask a person a question, with options when it has them (one recommended, several allowed when it says so); it waits up to the skill's human_wait_seconds (300 by default) with the job's timeout clock paused. The question opens a needs_human alert and shows the job in the inbox under Needs input, its options as buttons; list_inbox {view: needs_input}, list_jobs {waiting: true} and describe_cloud list the same.

answer_job {job, answer, option?, options?, resume?} answers it: answer is what the person said, option the option they picked, options the several they picked when the question allows more than one (they reach the agent one per line, before the answer). While the agent still waits, the answer is delivered live into the run. When the run already ended (the wait ran out, or the agent finished with outcome needs_human), resume: auto (the default) starts a new job with trigger resume that continues the Claude or Codex session with the answer; resume: never only records it. Answering resolves the alert. A run that ends with its question unanswered counts as needs_human.

Replaying a job or a delivery

  • replay_job {job}: the same payload again through the skill as it is now, as a new job with trigger replay. Nothing is deduplicated and the signature is not checked again.
  • replay_delivery {delivery, force?, skip_filters?, machine?, skill?}: a stored delivery again, including one the machine rejected or skipped. force is needed for a delivery that failed its checks (they are skipped on replay); skip_filters ignores the skill's when filters. With machine or skill, the stored body runs on another machine or through another skill as a fresh run: Webhook history and replay.

Replay once the cause is fixed: the runner signed in, the secret set, the skill corrected. Both need a fleet:run key or the member role, and control mode on the machine.

run_skill and test_skill

run_skill {machine, skill, payload?, headers?, runner?, model?, effort?, wait_seconds?} runs an installed skill as if a webhook arrived with that payload: no signature check, no filters. test_skill {machine, skill_md, payload?, runner?, wait_seconds?} runs a whole SKILL.md once without installing it (trigger test), to try a draft before save_skill; the dashboard's Playground does the same. Both answer with the command and, when asked to wait, the job: wait_seconds (up to 240) waits until the job finishes or asks a person, and the answer says pending when it is still on its way.

Live output

watch_job {job, ttl_seconds?, stream?} asks the machine to stream a running job's stdout (or stderr) to the cloud for ttl_seconds (900 by default, from 10 to 3,600); get_job and the job's page then show it as output, refreshed every few seconds. The protocol allows 5 watched jobs at a time.

Artifacts

A job is a directory on the machine. get_job_artifact {job, name} fetches one of its files: up to 256 KiB is returned (the end of stdout and stderr, the start of the others); larger files are uploaded by the machine in chunks and read from there. Secrets are scrubbed on the machine before anything is sent.

ArtifactTypeDescription
stdout
file
The agent's transcript: what it did, in the runner's own output format.
stderr
file
The runner's error output: where a crash or a login problem explains itself.
prompt
file
What the agent was told: the SKILL.md body with the placeholders filled, plus the guardrails. Quotes the webhook body.
result
file
The agent's final message.
payload
file
The webhook body the job got.
event
file
The whole event: body, redacted headers, query string and sender.
response
file
The agent's report: outcome, summary, links and data.

Cost and tokens

Each job records what the runner reported: the agent's cost in dollars and its tokens (Claude and Codex usage added up). The job's page and get_job show them; get_stats and the Stats page sum them per day and per skill, with p50 and p95 durations. A shell job has no cost.

Triggers

trigger says what started a job; list_jobs {trigger: ...} filters by it.

TriggerStarted by
webhookA delivery at the machine's own URL or through a hosted URL.
cliStarted by hand on the machine with skillhook run.
mcpStarted by an agent through the machine's own MCP server.
apiStarted through the machine's HTTP API.
scheduleA schedule: block came due; the payload says which slot.
replayreplay_job or replay_delivery: the same request again as a new job.
testtest_skill, or skillhook run --file: a SKILL.md that is not installed.
resumeAn answer to a question after the run had ended: a new job continues the agent's session.

Every one of these is a job the machine reports, so it appears in the dashboard and in list_jobs whether it was started from a webhook, the machine's terminal or the cloud. Jobs from a machine the cloud only observes appear too; acting on them needs control mode (see Machines and pairing).