Skip to content

Troubleshooting

When something breaks, the events page is your first stop — most failures surface there as a warning. Below are the issues operators hit most.

permission denied connecting to the Docker daemon socket

Section titled “permission denied connecting to the Docker daemon socket”

Cause. The agent’s group ID doesn’t match /var/run/docker.sock on the host, so it can’t talk to the Docker daemon.

Fix. On the host, either add the agent’s user to the host’s docker group, or loosen the socket’s group permissions:

Terminal window
sudo chmod g+rw /var/run/docker.sock

Sign-in succeeds but bounces back to /login

Section titled “Sign-in succeeds but bounces back to /login”

Cause. The auth origin doesn’t match the URL in your browser. Session cookies are scoped to that origin, so a mismatch means the cookie is never sent back.

Fix. Set RUNAWAY_HUB_URL to the exact user-facing https:// URL and restart the hub. If the browser reaches the hub at a different origin than the canonical URL, set RUNAWAY_APP_URL to that origin instead. See TLS and reverse proxy.

Cause. Usually a GitHub problem (the PAT is missing a scope or you’re rate-limited) or the workload has backed off after repeated failures.

Fix. Check the events page for github warnings and open the workload to see its backoff state. A backed-off workload retries on its own; fix the underlying GitHub error (see PAT issues below) and it recovers.

Jobs fail with Cannot find module 'node:path' (or similar Node-stdlib errors)

Section titled “Jobs fail with Cannot find module 'node:path' (or similar Node-stdlib errors)”

Cause. The cached runner image is stale — an old actions/runner that predates working Node externals.

Fix. Open the runner profile’s settings and set Image pull to Always or TtlHours. The next reconcile pass refreshes the layer and every spawn after runs on a current image. See Runner profiles and Registry credentials.

Private-registry pull fails with unauthorized

Section titled “Private-registry pull fails with unauthorized”

Cause. The runner image lives in a private registry and the hub has no credential for it.

Fix. Add an image credential and attach it to the runner profile’s image-pull credential select. See Registry credentials.

docker build / docker run fail inside a workload

Section titled “docker build / docker run fail inside a workload”

Cause. The workload’s runtime is none (the default), so runners have no Docker daemon to talk to.

Fix. Edit the workload’s runtime: shared-daemon shares the host’s Docker daemon, while dind and isolated-sysbox give each runner its own inner one. isolated-kata won’t help — it runs no daemon. Runtime changes apply to new spawns; in-flight runners finish on their current runtime.

A job can’t find files that are definitely in the repo

Section titled “A job can’t find files that are definitely in the repo”

Symptoms vary and none of them mention mounting: package.json not found, a CLI reported as missing, a corepack or nvm shim downloading a default version instead of the pinned one, a test run that collects zero files.

Cause. The job bind-mounted a path into a container on a shared-daemon workload, and the path doesn’t exist on the host. The host’s daemon resolved it, found nothing, created an empty directory and mounted that. The container starts normally and the failure surfaces further downstream.

Fix. Mount only paths under $GITHUB_WORKSPACE, which Runaway binds from the host at a matching path. See bind mounts. A nested tier doesn’t have this constraint at all.

container: job fails with exec: "/__e/node24/bin/node"

Section titled “container: job fails with exec: "/__e/node24/bin/node"”

Cause. The workload’s runtime is shared-daemon. The runner starts a container: job as a sibling on the host’s daemon and mounts the Node it injects for JavaScript actions from a path that exists only inside the runner image, so the job container gets an empty directory where the action runtime should be. services: fails the same way for a different reason: the container publishes its ports on the host, and the job looks for them on its own loopback.

Fix. Move the workload to dind or isolated-sysbox, where a job’s containers are children of the runner’s own daemon, or run the job on a GitHub-hosted runner. Neither pattern works on shared-daemon, and no path or network setting gets around it. See converting a workflow.

Disk fills with directories under /var/lib/runaway/work

Section titled “Disk fills with directories under /var/lib/runaway/work”

Cause. shared-daemon runners get a workspace tree on the host so the daemon can resolve a job’s bind mounts. AutoRemove takes the container but not the directory, and the agent can only reclaim the trees if /var/lib/runaway/work is bound into it read-write at that same path. An agent created before that mount existed keeps its original mount set across an image update, logs one notice, and leaves the trees alone.

Fix. Re-create the agent container from the current installer or compose file. Runners work either way — this costs disk, not correctness. See updating an agent.

network mode "…" not supported by buildkit

Section titled “network mode "…" not supported by buildkit”

Cause. A docker build passed --network a user-defined network. BuildKit accepts only host, none and default; joining another network needs a dedicated builder created with --driver-opt network=….

Fix. Publish a port on the service the build needs and reach it with --add-host=host.docker.internal:host-gateway. Note that --network host is not a workaround — BuildKit accepts the flag and ignores it, so the build stays in its own namespace and 127.0.0.1 resolves to the build container rather than the host.

An installer fails on an unexpected language version

Section titled “An installer fails on an unexpected language version”

For example pnpm/action-setup exiting from its self-installer, or a tool reporting a Node or Python version far older than the workflow pins.

Cause. The step ran before actions/setup-node / actions/setup-python, so it used the runner image’s system interpreter. On a GitHub-hosted runner that is recent enough not to matter; on the Standard image it is the base distribution’s.

Fix. Order toolchain actions before anything that shells out to them. See the Standard image.

Container exits immediately with a lock message

Section titled “Container exits immediately with a lock message”

Cause. Another container is holding the data-volume lock. SQLite is single-writer.

Fix. Run exactly one container per data volume. If you started a second one, stop it.

Cause. The agent isn’t connected — it’s stopped, the machine is down, or its token was revoked.

Fix. Restart the agent on the host and confirm it can reach the hub. If you revoked its token, mint a fresh installer URL and re-enroll (re-enrollment is non-destructive of the data volume). See Adding hosts.

A host won’t reconnect after the hub was reinstalled

Section titled “A host won’t reconnect after the hub was reinstalled”

Cause. The agent persists its identity — hub URL, agent token, and host id — to /data/agent.json inside the runaway-agent-data volume, and reuses it on every boot. If you reinstalled the hub or it lost its database, that saved token belongs to a hub that no longer exists, so the new hub can’t adopt the agent. Re-running the installer doesn’t help: the agent prefers the saved identity and ignores the fresh enrollment token while the volume is intact.

Fix. Wipe the stale identity on the host, then enroll fresh from the new hub. Remove the agent and its data volume:

Terminal window
docker rm -f runaway-agent
docker volume rm runaway-agent-data

If the old install left runner containers behind, clear them too — they carry the previous install’s labels and the new hub will never adopt them:

Terminal window
docker rm -fv $(docker ps -aq --filter "label=managed-by=runaway") 2>/dev/null || true

-v takes each runner’s anonymous volumes with it — on the dind tier that includes the inner daemon’s data root, which holds every image the runner pulled and is the largest thing it leaves. Named volumes are never removed this way, so organization caches survive.

Then add the host on the new hub and run the install command it gives you. See Adding hosts.

Cause. GitHub tokens expire or get edited, which surfaces as github warnings on the events page and stalls the affected org’s runners.

Fix. Re-enter the PAT for that org. See GitHub setup for the required scopes.

Report a problem with a diagnostics bundle

Section titled “Report a problem with a diagnostics bundle”

When a problem outlives the fixes above, download a diagnostics bundle and attach it to your issue instead of copying the events table by hand. Settings → Diagnostics → Download diagnostics, or use the Download diagnostics button on the events page.

The bundle is a single JSON file holding this install’s identity and versions, its hosts, workloads, and organizations, recent runners and jobs, and the last 24 hours of events — enough to reconstruct your setup and the failure timeline.

Secrets never enter it: credentials, registry passwords, and webhook secrets are left out entirely, customEnv values are masked, and the event log is scrubbed of token-shaped strings. It’s plain JSON, so you can read the whole file before sharing it.