Forgeplane workers and capability-based execution
A Forgeplane worker is the execution process. The coordinator owns authentication, policy, persistence, scheduling, and the run record; the worker prepares an isolated workspace, invokes an available infrastructure tool, and reports status, logs, artifacts, and state.
How scheduling works
Section titled “How scheduling works”- A worker registers with the coordinator and joins a named pool.
- It advertises capabilities and reports heartbeats over the worker gRPC API.
- The coordinator selects a connected worker that matches the required pool and tool capability.
- The coordinator records an assignment lease and publishes the job through NATS JetStream.
- The worker fetches the trusted execution bundle, runs the tool, and streams results.
Queue delivery does not replace assignment and lease checks. An enrollment token is required for initial registration. Use TLS, and mTLS where configured, outside local development.
Before invoking the tool, the worker claims the persisted execution attempt through the coordinator. It signs the safe terminal outcome, and the coordinator commits its receipt before acknowledging the JetStream message. Forgeplane never reassigns an attempt that crossed this execution fence merely because its outcome is missing. See Recover an unresolved execution.
Workers use gRPC to fetch execution bundles and managed state, stream logs and status, and upload artifacts or updated state.
Capabilities are an image contract
Section titled “Capabilities are an image contract”A worker executor in code does not prove that the corresponding binary is present in the container. The default production worker image intentionally bundles no IaC binaries and advertises an empty capability list:
[]A tool-enabled image must contain every binary it advertises and must advertise only those capabilities. Workers do not download IaC tools at runtime. An image with OpenTofu should advertise tofu, not terraform or ansible.
| Tool | Capability | Supported lifecycle |
|---|---|---|
| Terraform | terraform |
Plan/preview, apply, teardown, streamed logs, and managed state |
| OpenTofu | tofu (opentofu is an accepted alias) |
Plan/preview, apply, teardown, streamed logs, and managed state |
| Ansible | ansible |
Governed execution and streamed logs; no preview, teardown, managed state, or plan-based drift |
helm, pulumi, and bash exist as shared tool labels, but the current worker does not ship executors for them.
Worker pools
Section titled “Worker pools”The Helm chart models workers with shared workerDefaults and named workerPools. Defaults hold common runtime, probes, security, storage, and service settings. A pool can override its image, capabilities, workload kind, service, resources, and scaling.
workerDefaults: runtime: workspaceRoot: /var/lib/forgeplane/workspaces
workerPools: default: capabilities: []
tofu: image: repository: registry.example.com/forgeplane-worker-tofu tag: "reviewed-release" capabilities: - tofu extraEnv: - name: FORGEPLANE_WORKER_IAC_UID value: "10001" - name: FORGEPLANE_WORKER_IAC_GID value: "10001" - name: FORGEPLANE_WORKER_POOL_NAME value: tofuSet the registration pool explicitly; the chart fleet name alone does not set it. See Helm pool configuration for extraEnv replacement and execution-identity requirements.
Treat the image and capability list as one reviewed change. Worker workspaces are disposable and are not the canonical copy of managed state.
Capacity and lifecycle state
Section titled “Capacity and lifecycle state”Each worker reports a maximum concurrent-job count. The coordinator schedules only within the lower of that capacity and the worker_max_concurrent_jobs system setting, and stops assigning new work to draining or offline workers. Because a worker’s runs share one IaC UID, a worker currently reports a capacity of one and runs one job at a time; scale out by adding workers.
| Status | Meaning |
|---|---|
idle |
Connected and available |
busy |
At the current execution capacity |
draining |
Finishing work but accepting no new assignments |
offline |
No longer schedulable after missed heartbeats |
Drain a worker before maintenance. Reactivate or disconnect it from the admin worker view after the maintenance boundary is complete.
Retain a worker’s old outcome-signing public keys while attempts pinned to those keys remain recoverable. Key rotation must not make a late authoritative outcome unverifiable.
Execution and security boundary
Section titled “Execution and security boundary”The coordinator validates input bindings and materializes approved secret payloads into the execution bundle. Workers validate the bundle’s trust metadata before use.
Every Terraform, OpenTofu, and Ansible command runs as the isolated IaC UID/GID. A worker that cannot switch to it (one that is not running as root on Linux, or has no IaC UID configured) refuses IaC execution. A canceled or timed-out Terraform or OpenTofu command first receives SIGINT on its process group and has 60 seconds to save its state and exit. After that, and when any command finishes, the worker kills every remaining process of the IaC UID. It does the same before each job, failing that job if any process survives, and at startup, so no process from one run reaches the next run’s files or secrets.
Keep coordinator and worker credentials separate. Restrict worker egress to the providers it needs. Give worker pods writable space only for temporary directories and workspaces. Do not use a worker workspace as a backup or state authority.
The repository includes an opt-in local OpenTofu image for validation. It advertises only tofu; it does not add OpenTofu to the production worker image or change Helm defaults.
See Run Terraform and OpenTofu in private networks for deployment-pattern guidance, then review Docker Compose, Helm chart, and Managed state.