Machines and pairing
Observe and control modes, the policy a machine keeps, what happens offline, disconnecting.
Pairing codes
A machine joins an organisation with a pairing code from the dashboard's Pair a machine page (admins and owners). A code is shown once, works once, for ten minutes, and pairs one machine. The person making it chooses the mode; a code made for observe never yields control, whatever the machine asks for.
Pairing hands the machine a token of its own (SKILLHOOK_CLOUD_TOKEN in its .env; the cloud stores only a hash) and writes cloud.* to its skillhook.json. The cloud never holds the machine's admin token, and the token is never used for anything but the link: a paired machine cannot read the rest of its organisation. Pairing is done by a signed-in person; no API key can do it, so nothing a key does can outlive it. Pairing needs skillhook 0.5 or newer.
Observe and control
- Observe: read-only. Deliveries, jobs, progress, schedules, configuration and health changes flow to the cloud; the machine accepts only read commands (health, jobs, deliveries, stats, skills). Nothing acts on the machine from the cloud.
- Control (
--control): the cloud may also queue actions: run, test, replay and cancel jobs, answer an agent waiting for a person, patch the configuration, fire a schedule, write or remove skills, generate a secret, restart a service-run server, install an update.
fleet:run or fleet:admin key, can run code on that machine. Pair in observe mode where that is not acceptable, and narrow control mode with the machine's own policy.The machine's own policy
The cloud checks every command before queueing it (the caller's role, the arguments, the machine's state and the policy it last reported), and the machine checks again and has the last word. cloud.allow_commands and cloud.deny_commands in skillhook.json narrow what a controlled machine accepts; the cloud cannot widen them, and a command the machine refuses is shown as refused. Some settings are never the cloud's to change: host, port, trust_proxy, runners, env_passthrough, projects and everything under cloud.*. The machine also refuses to generate or set its own credentials (SKILLHOOK_ADMIN_TOKEN, SKILLHOOK_CLOUD_*) from the cloud, so the cloud cannot rotate the admin token or move the link.
The sync link
Machines pull. skillhook serve holds one outbound HTTPS request to the cloud open for up to 25 seconds; each request carries the machine's new events, command results and hosted-delivery acknowledgements, and each answer carries queued commands and hosted deliveries. Both directions are at-least-once: the machine re-sends until acknowledged, the cloud deduplicates on event and command ids. Events wait on the machine while the cloud is unreachable. The poll is the heartbeat: the cloud expects the next one within the hold plus a minute of grace, and a minute tick marks a machine offline once that passes. A machine whose link reports itself degraded shows as degraded, with the reason on its page; a sync that reports a healthy link makes it online again.
Health checks and runner readiness
A machine reports its health checks, grouped (system, skillhook, runners, tools, skills, exposure): the Claude Code and Codex versions and logins, every MCP server they know and whether it connects, free disk, the public URL, and per skill its last run and any env: name that is not set. The machine's page shows failing and warning checks with the fix each one suggests; a check that starts failing opens an alert, and a check that recovers resolves it. Runner readiness is checked as soon as the machine connects, so the dashboard says whether claude and codex are installed and signed in before the first job runs. A job whose runner is not ready fails at once with failure kind auth, unless the skill names a fallback runner that is ready.
What an offline machine means
An offline machine runs nothing. Commands queued for it wait up to 10 minutes and then expire; the dashboard and the API show each command's status, and nothing assumes a command ran. Hosted webhook URLs keep accepting deliveries and hold them, sealed, for 72 hours, until the machine collects them on its next sync. A machine_offline alert opens when the machine stops polling and resolves itself when it is back; providers such as Granola retry failed deliveries for days, so a short sleep loses nothing. Jobs run only while the machine is awake: on a desktop, prevent automatic sleep.
Rename and disconnect
- Rename (admins, or
rename_machinewith afleet:adminkey) changes the name the cloud shows; the hostname stays what the machine reports. - Disconnect (admins, or
disconnect_machine) revokes the machine's token and cancels its pending commands: it stops syncing at once. Its deliveries, jobs and history stay; pairing it again creates a new machine.skillhook cloud disconnecton the machine does the same from that side.
The protocol
The wire protocol is skillhook's, exported as @meterapp/skillhook/protocol, and this service imports it rather than re-declaring it. This deployment speaks protocol version 1; a machine that speaks a newer one is answered 426 upgrade_required until the service catches up (a unit test fails here when the installed skillhook speaks a version outside the range, so the range moves with each release). The messages, limits and commands are documented in skillhook's docs/cloud.md and docs/cloud-protocol.md.
Protocol commands by role
Every action on a machine is one of these commands, queued by the dashboard, by a tool such as run_skill, or by send_command with any of them. The lowest organisation role that may send each one is fixed in the cloud; an API key acts as the role of its scope (see API keys and scopes). Reads that can reveal configuration, logs or secret names need admin. The machine's mode and policy decide too.