Skip to main content
Version: v3.9.0

Process manager

The Process Manager reconciles desired workload state (from Controller, local manifests, and ControlPlane SQLite rows) with the active ContainerEngine. It runs a periodic monitor loop, a task queue for lifecycle operations, and publishes running counts to StatusReporter.

Code: internal/processmanager/

Purpose​

  • Pull/create/start/stop/remove containers for managed microservices
  • Reconcile local_workloads (CLI edgelet deploy) and system_control_plane deployments
  • Map container runtime state back to status structures for Field Agent POST
  • Queue lifecycle tasks (start/stop/restart/kill) from EdgeletAPI
  • Honor quiesce during engine restart (SetQuiesced)
  • Apply workload metadata labels/env via workloadmeta helpers

Dependencies​

Depends onReason
fieldagentMicroserviceManagerInterface: latest microservice list, registries
engine.ContainerEngineDocker, Podman, or edgelet/containerd CRI
storeLocal workloads, control plane row, runtime_container_refs
networkBridge/network setup for workloads
statusreporterProcessManager status, running counts
configEngine name, reconcile-cycle logging
Used byReason
supervisorStarted after engine wired
edgeletapi / runtimeapiMS lifecycle, deploy apply, logs, exec
pruningImage list callback from latest microservices
healthcheckEdgelet-engine healthcheck runner

Lifecycle​

Start​

(*ProcessManager).Start(engine, microserviceManager):

  1. Store engine reference and create ContainerManager
  2. Start goroutines: containersMonitor, containerStatsLoop, checkTasks
  3. Task queue capacity 100

Stop​

Cancel context, close task queue, wait for goroutines, drain shutdown with configurable timeout per container.

Reconcile loop​

containersMonitor() wakes about every 5 seconds, and immediately when a workload is marked. A pass reconciles only the workloads that are due.

A workload is reconciled when:

  • its spec changes
  • its container starts, exits, is OOM-killed, or is deleted
  • a backoff or volume wait is due
  • a catalog item it needs becomes ready or failed

A full compare of every workload still runs about once a minute.

With the embedded engine and a healthy event stream, an idle workload is not inspected every 5 seconds. If the event stream is down, inspection returns to every 5 seconds until the stream is healthy. Docker and Podman still check running-or-not every 5 seconds; a container that is not running is reconciled on that check. CPU and memory on status refresh about every 10 seconds, on containerStatsLoop(), separate from reconcile.

When a pass has work, due workloads run in this order:

reconcileControlPlane() // when that workload is due
reconcileControllerMicroservices() // marked controller UUIDs
reconcileMarkedLocal() // marked local_workloads
deleteRemainingMicroservices()
pruneStaleProcessManagerStatuses()
updateRunningMicroservicesCount()
updateCurrentMicroservices()

An idle pass with a healthy event stream does not load containers. Reconcile is skipped while IsQuiesced() (pending engine restart). When the engine is ready again, the next pass is a full compare.

Startup and reconnect call Update(), which marks every workload once and wakes the monitor.

Workload sources​

SourceSQLiteEdgeletAPI source filter
Controllercontroller_microservicesmanaged
Local deploylocal_workloadslocal
ControlPlanesystem_control_planecontrolplane

Each running workload may have a row in runtime_container_refs linking microservice UUID, scope (controller | local), workload ID, and sandbox ID.

Local and ControlPlane reconcile​

Local and ControlPlane deployments use desired-state fields (desired_state, runtime_state, generation, observed_generation, failure_count):

  • running. Ensure container exists and matches manifest generation
  • stopped. Stop container, keep record
  • deleted. Remove container and delete the local_workloads row (no persistent tombstone). Reconcile re-reads the row before write and never inserts a missing UUID.

A local recreate is stored as starting, with the new generation not yet observed, before the old container is removed. A delete of that container does not start another container or treat the new generation as observed while that apply is still in progress.

When the ControlPlane workload is due, it is reconciled before managed microservices so the controller container is stable before dependent workloads.

See Workload continuity for restart and engine-switch behavior.

Configuration​

KeyEffect
containerEngineRecorded in runtime state; affects healthcheck path and whether reconcile uses the event stream
logReconcileCycleEveryNTicksStructured reconcile cycle logging

The 5-second wake, the full compare (about once a minute), and the CPU and memory sample (about every 10 seconds) are built in. config.yaml has no keys for them.

External APIs​

Process Manager does not serve HTTP directly. EdgeletAPI routes delegate through internal/runtimeapi:

OperationEdgeletAPI examples
List/inspect MSGET /v1/ms, GET /v1/ms/{id}
LifecyclePOST /v1/ms/{id}/start, POST /v1/ms/{id}/stop, POST /v1/ms/{id}/restart, POST /v1/ms/{id}/kill
Local deployPOST /v1/deploy/microservices:apply
Logs/execGET /v1/ms/{id}/logs, exec session routes

Observability​

  • Log module name: "Process Manager"
  • StatusReporter index: 1 (utils.ProcessManager)
  • ProcessManagerStatus: per-microservice status map, registry status, running count
  • Reconcile cycle events when logging enabled (runtimeops emit helpers)

Debug codes: PMCM (containers monitor), PMCT (check tasks).

Status, crashes, and restarts​

Process Manager publishes per-microservice status to StatusReporter. Field Agent and edgelet ms inspect read the same fields.

FieldBehavior
errorMessageCurrent failure. Kept while the workload is failing, restarting, or has been RUNNING for less than 30 seconds after a crash. After 30 seconds of continuous RUNNING, the current field is cleared to ""
lastError / lastErrorAtLast crash text and unix-ms timestamp. Overwritten on a new failure. Not cleared on recovery or rebuild
restartCountReal restart events (RUNNING→EXITING, failed start, or a scheduled crash-driven recreate). Omit when 0. Reset to 0 on operator rebuild only

Crash text:

EngineExample
Docker / PodmanexitCode=1 oomKilled=false error=config missing
Embedded (edgelet)CRI reason=... exitCode=... message=...

Crash loops delay recreate (10s, 20s, ... up to 5 minutes). Status stays the real runtime state (EXITING, CREATED, QUEUED if the container is gone). There is no extra state name for “waiting to recreate”. Operator rebuild, catalog becoming Ready, and a single non-restartable CRI recreate skip the delay.

STUCK_IN_RESTART means 10 real restarts in 10 minutes, not monitor ticks. Operator rebuild retries. After five exhausted lifecycle tasks the workload stays FAILED until rebuild (existing skip-until-rebuild).

Last crash text for controller-managed workloads is in-memory. An agent restart may drop it until the next failure. Local and ControlPlane rows keep existing SQLite last_error / restart_count (not cleared on first start).

Failure modes​

SymptomTypical cause
MS stuck updatingReconcile quiesced; engine unavailable
Crash text missing after agent restartController-managed last crash is in-memory; wait for the next exit or inspect
Recreate not immediate after a crashCrash-loop backoff (10s ... 5 min); status is still EXITING (or similar), not a new state
STUCK_IN_RESTARTTen real restarts in ten minutes; use operator rebuild
Local deploy failuresfailure_count threshold; image pull errors
CP container recreatedRow still in system_control_plane; external docker rm
ms rm on CP rejectedBy design: use edgelet controlplane delete

Code map​

FileRole
manager.goMonitor loop, Update, task queue, crash-loop backoff
reconcile_schedule.goWhich workloads are due on a pass
container_stats.goCPU and memory sample loop
container_manager.goEngine CRUD for containers
status_sync.goCurrent vs last error, 30s RUNNING grace
restart_checker.goReal restart events → STUCK_IN_RESTART
controlplane_reconcile.goControlPlane desired-state machine
local_launch.goLocal manifest launch paths
controlplane_ops.goCP-specific engine operations
quiesce.goEngine restart quiesce flag
lifecycle_*.goStart/stop/restart implementations

Related: Field agent, Store, Container engines, Troubleshooting.

Group 3See anything wrong with the document? Help us improve it!