Troubleshooting
Common issues when running Edgelet on edge nodes.
Daemon won't start
Symptoms: systemctl start edgelet fails; no listening socket on :54321.
Checks:
-
Validate config paths and permissions:
ls -la /etc/edgelet/edgelet system info -
Review journal logs:
sudo journalctl -u edgelet -n 100 --no-pagertail -f /var/log/edgelet/daemon-startup.log -
Confirm
containerEngineis valid on this platform:edgelet system version -o json | jq '{allowedEngines, containerEngine}'grep containerEngine /etc/edgelet/config.yamlLinux allows
edgelet,docker, orpodman. Darwin/windows allowdockerorpodmanonly. -
Check disk space:
df -h /var/lib/edgelet /var/lib/edgelet-containerdIf usage grows after deleting microservices, private
VOLUMEdata remains under/var/lib/edgelet/volumes/data/and shared names under/var/lib/edgelet/volumes/shared/until you reclaim withedgelet volume. See Volumes. Scheduled prune andedgelet system prunedo not delete those trees.
Cgroup delegation (embedded edgelet engine)
Symptoms: daemon fails at startup with controller cpu is not available; RunPodSandbox errors; nested Docker deploy fails immediately.
Checks:
-
Confirm cgroup mode and driver:
edgelet system status -o json | jq '{cgroupMode,cgroupDriver,cgroupNested,cgroupDelegatedControllers}' -
On cgroup v2 hosts, verify delegated controllers in the edgelet cgroup:
grep -E '^(cpu|memory|pids)' /sys/fs/cgroup/cgroup.controllerscat /sys/fs/cgroup/cgroup.subtree_control -
Hybrid v1+v2: edgelet logs a warning and prefers unified v2. Migrate the host to pure cgroup v2 when possible. See cgroups.
-
Nested edgelet container in Docker:
--privilegedis required:
docker run -d --name edgelet --privileged \
-v /var/lib/edgelet:/var/lib/edgelet \
-v /etc/edgelet:/etc/edgelet \
ghcr.io/datasance/edgelet:<tag>
Without --privileged, cpu/memory/pids controllers are not delegated and CRI cannot create sandboxes.
-
Non-systemd init (openrc, sysvinit, s6): edgelet uses the cgroupfs driver automatically. Ensure
/sys/fs/cgroupis writable and controllers are mounted. -
LXC/VM machine root (
/.lxc, OpenRC on OrbStack Alpine): error mentionsedgelet-cgroup-prepor empty/.lxc/cgroup.controllersafter cold boot:rc-status | grep edgelet-cgroup-prepcat /proc/self/cgroupcat /sys/fs/cgroup/.lxc/cgroup.controllersReinstall edgelet so
install.shregistersedgelet-cgroup-prepin sysinit, then reboot:sudo ./install.sh --version=<tag>sudo rebootAfter reboot,
edgelet cgroup-preflightshould pass andrc-service edgelet-containerd startshould succeed without manual cgroup commands.If
edgelet-containerdfails after a prior successful start withpreparing machine-root cgroup delegationor immediate spawn failure in/var/log/edgelet/containerd.log, upgrade to a current beta.2 build (skips reparent when prep already ran and prepares OpenRC staging +/edgeletcgroups). Do not run manual Moby reparent commands ifedgelet-cgroup-prepalready ran at sysinit. -
RunPodSandbox/sd-bus call: Invalid unit name or type: crun withSystemdCgroup=true(manualconfig.doverride or an old edgelet build). Reinstall the current edgelet build. Generated config must haveSystemdCgroup = falsefor crun; systemd bare-metal hosts must stay inedgelet.servicewith nocgroup.path. Verify:edgelet system status -o json | jq '{cgroupDriver,cgroupContainerdPath}'grep -E '^\[cgroup\]|path =|SystemdCgroup' /var/lib/edgelet-containerd/config.tomlcat /proc/$(systemctl show edgelet -p MainPID --value)/cgroup -
bpf pin to /sys/fs/bpf/crun/k8s_io/...: often appears whenSystemdCgroup=truewas enabled for crun. Use a current edgelet build (SystemdCgroup = false). If BPF is still required on your host, ensure/sys/fs/bpfis mounted and/sys/fs/bpf/crun/k8s_ioexists (see cgroups).
See cgroups for the full matrix.
Control vs data plane restart
Symptoms: MS disappear after systemctl restart edgelet; unexpected downtime during agent OTA.
Checks:
-
Identify which unit you restarted:
systemctl status edgelet edgelet-containerd --no-pager -
docker/podman: control restart should not drain MS (
shutdownPolicy=leave-runningdefault):grep shutdownPolicy /etc/edgelet/config.yamledgelet system status -o json | jq '."runtime.shutdownPolicy", ."runtime.agentPhase"'docker ps --filter label=iofog.org/microservice-uuid -
embedded split: restart
edgeletonly for control OTA; restartedgelet-containerdonly when the fat/runtime bundle changed (expect MS stop + reconcile). -
Monolithic embedded (no
edgelet-containerdactive):restart edgeletstill drains MS. Enable runtime split per Workload continuity. -
Cold engine change: changing
containerEnginealways recreates MS. This is expected and unrelated to workload continuity. -
Legacy unit on VM: if
systemctl cat edgelet-containerdshowsPartOf=edgelet.service, reinstall packaging and runsystemctl daemon-reload. That forces data-plane stop on control restart (embedded control-only restart gate failure).
See Workload continuity for the full OTA matrix.
Embed bundle / data-plane restart
Symptoms: edgelet-containerd in a crash loop; journal shows rename extracted bundle: file exists; Preparing data dir on every start; MS stuck after repeated data-plane restarts.
Diagnostic:
sudo journalctl -u edgelet-containerd -n 80 --no-pager
readlink /var/lib/edgelet/data/current/bin/aux/iptables # expect xtables-legacy-multi
ls -la /var/lib/edgelet/data/current /var/lib/edgelet/data/.lock
Recovery ladder:
-
Stop both units and clear systemd start-limit (if applicable):
sudo systemctl stop edgelet edgelet-containerdsudo systemctl reset-failed edgelet-containerdOpenRC:
rc-service edgelet stop && rc-service edgelet-containerd stop -
In-place repair (aux symlink drift):
sudo ln -sf xtables-legacy-multi /var/lib/edgelet/data/current/bin/aux/iptables -
Start data plane; wait for readiness:
sudo systemctl start edgelet-containerdsudo journalctl -u edgelet-containerd -f # Embedded containerd is readyOpenRC:
rc-service edgelet-containerd start -
Start control plane:
sudo systemctl start edgelet -
Nuclear (last resort. Re-extracts bundle; MS stop + reconcile):
sudo systemctl stop edgelet edgelet-containerdsudo rm -rf /var/lib/edgelet/data/<hash-dir> /var/lib/edgelet/data/<hash-dir>-tmp /var/lib/edgelet/data/.locksudo systemctl start edgelet-containerdsleep 5sudo systemctl start edgeletReplace
<hash-dir>with the bundle directory name under/var/lib/edgelet/data/(notcurrentorprevioussymlinks). After a healthy daemon start, only those two hash trees remain; leftover extracts from older upgrades are removed automatically.
Prefer systemctl stop then systemctl start over blind restart during shim upgrades. See Container engines. After five rapid failures within 300s, systemd stops auto-restarting edgelet-containerd until reset-failed (openrc: respawn_max=5 per 300s window).
Upgrade stays on the old version
Symptoms: An upgrade finishes without changing the installed version. The data-plane journal or install output contains exec data-plane drain: permission denied. findmnt /run shows noexec. Releases that stage the drain runtime under /run/edgelet/runtime-drain/ cannot execute it on that mount, so a fat upgrade stops before the thin binary is replaced.
Checks:
findmnt /run
sudo tail -n 200 /var/log/edgelet/ota-install.log
ls -l /var/lib/edgelet/data/.runtime-drain/edgelet
/usr/local/bin/edgelet version --verbose
readlink /var/lib/edgelet/data/current
/var/log/edgelet/ota-install.log is the detached install.sh output from a controller upgrade. Each attempt truncates the file, then writes that attempt. A non-zero install.sh exit is also written to the edgelet service log.
Recovery: Install a release that stages the drain runtime under diskDirectory (/var/lib/edgelet/data/.runtime-drain/ by default). /run can remain noexec. Then run the upgrade again. If the log says the drain did not verify, see Leftover process holding a volume.
Catalog runtime / orphan shims after data-plane restart
Symptoms: device or resource busy on sandbox or task cleanup; dial unix:///run/edgelet/containerd.sock: timeout; journal mentions left-over process or containerd-shim-* after edgelet-containerd stop; catalog workloads (WASM, spin, etc.) fail to start until manual cleanup.
Checks:
sudo journalctl -u edgelet-containerd -n 80 --no-pager
pgrep -af 'containerd-shim-.*edgelet/containerd.sock' || true
pgrep -af -- '--edgelet-containerd-child' || true
Recovery:
-
Stop both units:
sudo systemctl stop edgelet edgelet-containerdOpenRC:
rc-service edgelet stop && rc-service edgelet-containerd stop -
Reap orphaned data-plane processes:
sudo edgelet runtime reap-orphans -
Clear systemd start-limit if the data plane was crash-looping:
sudo systemctl reset-failed edgelet-containerd -
Start data plane, then control:
sudo systemctl start edgelet-containerdsleep 3sudo systemctl start edgelet
The command only targets edgelet embedded containerd (--edgelet-containerd-child and shims referencing /run/edgelet/containerd.sock). It does not kill host Docker or Podman shims.
Prefer stop then start on the data plane when upgrading catalog shims. See Container engines.
Containerd socket (edgelet engine)
Symptoms: connection refused to /run/edgelet/containerd.sock; microservices stuck in pull/create.
Checks:
-
Verify socket exists:
ls -la /run/edgelet/containerd.sock -
Check embedded containerd logs in daemon journal output.
-
Shim load errors such as
address: no such fileusually mean stale runtime task metadata. Restart the data plane first (systemctl restart edgelet-containerd); stale task cleanup runs on the next data-plane bootstrap. Then restart control if needed:systemctl restart edgelet-containerd.servicesleep 3systemctl restart edgelet.service -
Look for orphan overlay mounts blocking cleanup:
mount | grep edgelet -
Restart daemon (graceful):
sudo systemctl restart edgelet
EdgeletAPI 401 / auth failures
Symptoms: CLI exits 3 (Unauthorized); curl returns 401.
Checks:
-
Token file present and readable:
sudo ls -la /etc/edgelet/edgelet-api -
CLI uses correct CA:
edgelet auth whoamiedgelet auth whoami -o json -
After provisioning, ensure you are not sending an unsigned bootstrap JWT.
-
Regenerate PKI only when directed. Mismatched CA breaks existing CLI sessions.
Controller status 401 / unexpected deprovision
Symptoms: Agent deprovisions during controller restart or OTA; logs show repeated PUT status 401 or 503; controller-host node loses controller SQLite.
Context (v1.0.2+): Edgelet no longer deprovisions on the first status 401. Auto-deprovision runs only after 5 consecutive status auth failures spanning at least 60 seconds, and is suppressed while OTA reprovision is pending or Provision() is in flight. Status 503 and retryable: true responses never deprovision.
Checks:
-
Controller version. Full structured errors and readiness
/statussemantics require Controller ≥ v3.8.2. Older controllers still benefit from the streak gate but may return legacy error bodies. -
During OTA or controller restart, expect transient 401 or 503 on status POST and ping; the agent should stay provisioned and retry.
-
If deprovision still occurs after sustained auth failures, verify the provision key was revoked (
delete-nodestill deprovisions immediately) or Edge Guard hardware drift (see EdgeGuard). -
Inspect field agent logs for
status auth failuredeferral messages vs final deprovision after the gate opens.
CLI connectivity (exit 10)
Symptoms: Error[DAEMON_UNAVAILABLE]; exit code 10.
Checks:
-
Daemon running:
systemctl is-active edgeletedgelet system status -
Socket override. If using
--socket, path must match/run/edgelet/edgelet.sock. -
Firewall. EdgeletAPI listens on localhost/TLS; remote access is not the default operator path.
External Docker / Podman
Symptoms: Cannot connect to Docker daemon; containers not starting.
sudo systemctl status docker # or podman.socket
ls -la /var/run/docker.sock
grep containerEngineUrl /etc/edgelet/config.yaml
Ensure the configured containerEngineUrl matches the running engine socket.
Leftover process holding a volume
Symptoms: microservice status or errorMessage is volume in use by a leftover process; a new container logs Cannot lock file or another exclusive-lock error and exits (databases that exclusive-lock a file often do this). Common after a fat OTA or a data-plane stop whose drain did not verify. The volume files are still on disk. A host process still has them open.
Do not run edgelet volume rm (or delete volumes/data/ / volumes/shared/) to clear the lock. That destroys retained data and does not stop the process that holds the file.
Checks:
-
Confirm the wait text:
edgelet ms inspect <uuid|namespace.name> --summary -
Read the data-plane journal and the upgrade log. Failed drain is
module=RUNTIME_BOOTSTRAPwithdrain_verify_failed,drain_timeout, ordrain_degraded, anddata-plane drain did not verify.install.shprintsData-plane drain did not verify; binary was not replacedand leaves the previous thin binary installed.sudo journalctl -u edgelet-containerd -n 200 --no-pager | grep RUNTIME_BOOTSTRAP -
Find the process that still has the volume open (default
diskDirectoryis/var/lib/edgelet):sudo lsof +D /var/lib/edgelet/volumes/data /var/lib/edgelet/volumes/sharedsudo fuser -v /var/lib/edgelet/volumes/data /var/lib/edgelet/volumes/shared
Recovery: stop the process that holds the volume. Reconcile keeps the current runtime state and the text volume in use by a leftover process until that process is gone, including after an operator rebuild. The next cycle creates the container once the path is free.
On a node that is already up, stopping the holder is enough. Do not restart the data plane unless you intend to stop every workload.
When an upgrade or stop edgelet-containerd aborted, the journal line is leaving containerd running. Drain did not verify: runtime-bootstrap cleared the drain hold and stayed running with containerd still up (the unit did not exit, so systemd did not restart it). A leftover workload process may still hold the volume. KillMode=process means stop only signaled the parent; the containerd child can still be serving CRI. edgelet runtime reap-orphans does nothing until /run/edgelet/drain-verified exists. Stop the holder, then drain again while the CRI socket is up:
sudo edgelet runtime drain --direct
Exit 0 means verify passed. Re-run install.sh for an aborted upgrade, or stop and start the data plane only after that:
sudo systemctl stop edgelet-containerd
sudo systemctl start edgelet-containerd
OpenRC: rc-service edgelet-containerd stop, then start, after drain --direct exits 0.
Microservice crash / restart loop
Symptoms: dashboard errorMessage empty during a crash; Docker/Podman exit looks silent; recreate hammers the node; STUCK_IN_RESTART after a container sat in EXITING. Cannot lock file on a persistent volume is a leftover process, not this crash loop. See Leftover process holding a volume.
Checks:
-
Inspect the workload. Crash fields are the same on
GET /v1/ms/{id}andedgelet ms inspect:edgelet ms inspect <uuid|namespace.name> --summaryerrorMessageis the current failure. It stays set while the workload is failing, restarting, or has been RUNNING for less than 30 seconds after a crash. After 30 seconds of continuous RUNNING it is sent as"".lastError/lastErrorAtis the last crash text (unix ms). It is not cleared on recovery. Operator rebuild resetsrestartCountto 0 and keepslastError.- Docker/Podman text looks like
exitCode=N oomKilled=…(pluserror=…when the engine error is set). The embedded engine keepsCRI reason=….
-
Crash loops delay recreate (10s, 20s, ... up to 5 minutes). Status stays the real runtime state (
EXITING, and so on). There is no extra “waiting” state. Operator rebuild, catalog becoming Ready, and a single non-restartable CRI recreate skip the delay. -
STUCK_IN_RESTARTmeans 10 real restarts in 10 minutes, not status-poll ticks. Use operator rebuild to retry. -
Last crash text for controller-managed workloads is in-memory. Restarting
edgeletmay drop it until the next failure.
Controller connectivity
Symptoms: connectionToController not ok in edgelet system status.
edgelet system status -o json | jq .connectionToController
curl -v "$(grep ^controller: /etc/edgelet/config.yaml | awk '{print $2}')/health"
edgelet config cert <base64-controller-cert> # if TLS verification fails
Embedded integration tests
For embedded-engine regressions on macOS, use the Lima VM pipeline:
./test/embedded/run-all.sh
See ../../test/embedded/README.md.
Microservice exec
Symptoms: Error[EXEC_START_TIMEOUT]; attach fails with Error[NOT_FOUND]; controller exec works but local ms exec fails (or vice versa).
Checks:
-
Confirm the microservice container is running:
edgelet ms inspect <uuid|namespace.name>edgelet ms ls -
Retry after a start timeout. POST waits up to 15 seconds for the shell:
edgelet ms exec <uuid> -- /bin/shSee Exec sessions for the multi-session model and local vs controller limits.
-
Concurrent sessions: local CLI exec is unlimited per microservice; controller exec is capped at 3 per microservice on Controller. Multiple local sessions should not block each other after v1.0.0-rc.6.
-
Orphan containerd exec (embedded engine): if a prior exec crashed without cleanup, retry after the process manager orphan sweep or restart the microservice:
edgelet ms restart <uuid> -
Controller exec only: verify agent connectivity and that Controller v3.8.x (multi-exec sessions) is deployed. Status should list active controller session ids:
edgelet system status -o json | jq '.connectionToController' -
execEnabled: edgelet no longer gates exec on this flag. Session poll drives controller exec. Do not expect togglingexecEnabledto fix attach issues.
Collecting diagnostics
sudo journalctl -u edgelet --since "1 hour ago" > edgelet-journal.log
edgelet system status -o json > edgelet-status.json
edgelet system info -o json > edgelet-info.json
edgelet system version -o json > edgelet-version.json
Do not share /etc/edgelet/edgelet-api or private keys in support tickets.