Troubleshooting
When something breaks, the events page is your first stop — most failures surface there as a warning. Below are the issues operators hit most.
permission denied connecting to the Docker daemon socket
Section titled “permission denied connecting to the Docker daemon socket”Cause. The agent’s group ID doesn’t match /var/run/docker.sock on the host, so it can’t talk to
the Docker daemon.
Fix. On the host, either add the agent’s user to the host’s docker group, or loosen the socket’s
group permissions:
sudo chmod g+rw /var/run/docker.sockSign-in succeeds but bounces back to /login
Section titled “Sign-in succeeds but bounces back to /login”Cause. The auth origin doesn’t match the URL in your browser. Session cookies are scoped to that origin, so a mismatch means the cookie is never sent back.
Fix. Set RUNAWAY_HUB_URL to the exact user-facing https:// URL and restart the hub. If the
browser reaches the hub at a different origin than the canonical URL, set RUNAWAY_APP_URL to that
origin instead. See TLS and reverse proxy.
Runner containers don’t appear
Section titled “Runner containers don’t appear”Cause. Usually a GitHub problem (the PAT is missing a scope or you’re rate-limited) or the workload has backed off after repeated failures.
Fix. Check the events page for github warnings and open the
workload to see its backoff state. A backed-off workload retries on its own; fix the underlying
GitHub error (see PAT issues below) and it recovers.
Jobs fail with Cannot find module 'node:path' (or similar Node-stdlib errors)
Section titled “Jobs fail with Cannot find module 'node:path' (or similar Node-stdlib errors)”Cause. The cached runner image is stale — an old actions/runner that predates working Node
externals.
Fix. Open the runner profile’s settings and set Image pull to Always or TtlHours. The next
reconcile pass refreshes the layer and every spawn after runs on a current image. See
Runner profiles and
Registry credentials.
Private-registry pull fails with unauthorized
Section titled “Private-registry pull fails with unauthorized”Cause. The runner image lives in a private registry and the hub has no credential for it.
Fix. Add an image credential and attach it to the runner profile’s image-pull credential select. See Registry credentials.
docker build / docker run fail inside a workload
Section titled “docker build / docker run fail inside a workload”Cause. The workload’s runtime is none (the default), so runners have no Docker daemon to talk
to.
Fix. Edit the workload’s runtime: shared-daemon shares the
host’s Docker daemon, while dind and isolated-sysbox give each runner its own inner one.
isolated-kata won’t help — it runs no daemon. Runtime changes apply to new spawns; in-flight
runners finish on their current runtime.
A job can’t find files that are definitely in the repo
Section titled “A job can’t find files that are definitely in the repo”Symptoms vary and none of them mention mounting: package.json not found, a CLI reported as
missing, a corepack or nvm shim downloading a default version instead of the pinned one, a test run
that collects zero files.
Cause. The job bind-mounted a path into a container on a shared-daemon workload, and the path
doesn’t exist on the host. The host’s daemon resolved it, found nothing, created an empty directory
and mounted that. The container starts normally and the failure surfaces further downstream.
Fix. Mount only paths under $GITHUB_WORKSPACE, which Runaway binds from the host at a matching
path. See bind mounts. A nested tier
doesn’t have this constraint at all.
container: job fails with exec: "/__e/node24/bin/node"
Section titled “container: job fails with exec: "/__e/node24/bin/node"”Cause. The workload’s runtime is shared-daemon. The runner starts a container: job as a
sibling on the host’s daemon and mounts the Node it injects for JavaScript actions from a path that
exists only inside the runner image, so the job container gets an empty directory where the action
runtime should be. services: fails the same way for a different reason: the container publishes
its ports on the host, and the job looks for them on its own loopback.
Fix. Move the workload to dind or isolated-sysbox, where a job’s containers are children of
the runner’s own daemon, or run the job on a GitHub-hosted runner. Neither pattern works on
shared-daemon, and no path or network setting gets around it. See
converting a workflow.
Disk fills with directories under /var/lib/runaway/work
Section titled “Disk fills with directories under /var/lib/runaway/work”Cause. shared-daemon runners get a workspace tree on the host so the daemon can resolve a job’s
bind mounts. AutoRemove takes the container but not the directory, and the agent can only reclaim
the trees if /var/lib/runaway/work is bound into it read-write at that same path. An agent created
before that mount existed keeps its original mount set across an image update, logs one notice, and
leaves the trees alone.
Fix. Re-create the agent container from the current installer or compose file. Runners work either way — this costs disk, not correctness. See updating an agent.
network mode "…" not supported by buildkit
Section titled “network mode "…" not supported by buildkit”Cause. A docker build passed --network a user-defined network. BuildKit accepts only host,
none and default; joining another network needs a dedicated builder created with
--driver-opt network=….
Fix. Publish a port on the service the build needs and reach it with
--add-host=host.docker.internal:host-gateway. Note that --network host is not a workaround —
BuildKit accepts the flag and ignores it, so the build stays in its own namespace and 127.0.0.1
resolves to the build container rather than the host.
An installer fails on an unexpected language version
Section titled “An installer fails on an unexpected language version”For example pnpm/action-setup exiting from its self-installer, or a tool reporting a Node or
Python version far older than the workflow pins.
Cause. The step ran before actions/setup-node / actions/setup-python, so it used the runner
image’s system interpreter. On a GitHub-hosted runner that is recent enough not to matter; on the
Standard image it is the base distribution’s.
Fix. Order toolchain actions before anything that shells out to them. See the Standard image.
Container exits immediately with a lock message
Section titled “Container exits immediately with a lock message”Cause. Another container is holding the data-volume lock. SQLite is single-writer.
Fix. Run exactly one container per data volume. If you started a second one, stop it.
A host shows offline
Section titled “A host shows offline”Cause. The agent isn’t connected — it’s stopped, the machine is down, or its token was revoked.
Fix. Restart the agent on the host and confirm it can reach the hub. If you revoked its token, mint a fresh installer URL and re-enroll (re-enrollment is non-destructive of the data volume). See Adding hosts.
A host won’t reconnect after the hub was reinstalled
Section titled “A host won’t reconnect after the hub was reinstalled”Cause. The agent persists its identity — hub URL, agent token, and host id — to /data/agent.json
inside the runaway-agent-data volume, and reuses it on every boot. If you reinstalled the hub or it
lost its database, that saved token belongs to a hub that no longer exists, so the new hub can’t adopt
the agent. Re-running the installer doesn’t help: the agent prefers the saved identity and ignores the
fresh enrollment token while the volume is intact.
Fix. Wipe the stale identity on the host, then enroll fresh from the new hub. Remove the agent and its data volume:
docker rm -f runaway-agentdocker volume rm runaway-agent-dataIf the old install left runner containers behind, clear them too — they carry the previous install’s labels and the new hub will never adopt them:
docker rm -fv $(docker ps -aq --filter "label=managed-by=runaway") 2>/dev/null || true-v takes each runner’s anonymous volumes with it — on the dind tier that includes the inner
daemon’s data root, which holds every image the runner pulled and is the largest thing it leaves.
Named volumes are never removed this way, so organization caches survive.
Then add the host on the new hub and run the install command it gives you. See Adding hosts.
A PAT expired or lost scopes
Section titled “A PAT expired or lost scopes”Cause. GitHub tokens expire or get edited, which surfaces as github warnings on the events page
and stalls the affected org’s runners.
Fix. Re-enter the PAT for that org. See GitHub setup for the required scopes.
Report a problem with a diagnostics bundle
Section titled “Report a problem with a diagnostics bundle”When a problem outlives the fixes above, download a diagnostics bundle and attach it to your issue instead of copying the events table by hand. Settings → Diagnostics → Download diagnostics, or use the Download diagnostics button on the events page.
The bundle is a single JSON file holding this install’s identity and versions, its hosts, workloads, and organizations, recent runners and jobs, and the last 24 hours of events — enough to reconstruct your setup and the failure timeline.
Secrets never enter it: credentials, registry passwords, and webhook secrets are left out entirely,
customEnv values are masked, and the event log is scrubbed of token-shaped strings. It’s plain JSON,
so you can read the whole file before sharing it.