Skip to content

Forgeplane Helm deployment on Kubernetes

The Forgeplane chart deploys the coordinator, database initialization, and one or more named worker pools. It can install development-sized dependencies, or connect to services you operate separately. Production deployments must make explicit choices for durability, TLS, artifact storage, managed state, and backups.

You need:

  • Kubernetes and Helm 3;
  • signed coordinator, worker, and migrator images for the release;
  • PostgreSQL, NATS JetStream, and local or S3-compatible artifact storage;
  • cert-manager CRDs when certManager.enabled=true; and
  • a runtime Secret containing the keys required by enabled database, queue, storage, bootstrap, mail, and metrics modes.

Build chart dependencies before installing from a source checkout:

Terminal window
make chart-dependencies
helm install forgeplane charts/forgeplane \
--namespace forgeplane \
--create-namespace \
-f values-prod.yaml

Production secret defaults are intentionally empty. A raw render without suitable values should fail rather than emit placeholder credentials.

Maintainers can explicitly publish a development build from main without a Git tag or GitHub release. Each workflow attempt gets a unique 0.0.0-dev.<run-id>.<attempt> version and a successful build receipt recording the source commit, image digests, and chart version/OCI digest. The publisher does not update latest or deploy anything.

Follow the repository’s development-build runbook for publication approval, signature verification, private registry access, and reviewed GitOps pins. Use one complete successful receipt, not artifacts from failed or mixed attempts; deployment approval remains separate.

The chart can install these dependencies or connect to external services:

Service Subchart Enable value Production choice
PostgreSQL bitnami/postgresql postgresql.enabled Prefer a managed or independently operated database with tested backups.
NATS JetStream nats/nats nats.enabled Configure authentication, TLS, persistence, and clustering deliberately.
RustFS rustfs/rustfs rustfs.enabled S3-compatible option for artifact and state blobs.

coordinator.existingSecret replaces only the Forgeplane runtime Secret. It does not supply credentials required by enabled PostgreSQL, NATS, or RustFS subcharts.

For external PostgreSQL, use postgresql.connection.sslMode=verify-full and a hostname that matches the server certificate. The coordinator and migrator use the image trust store; the chart does not add a private database CA or client certificate automatically.

Every external NATS server must use max_payload >= 4194304 so the existing 2 MiB structured-output contract and its signed execution-outcome envelope fit. If Forgeplane reports the observed limit, raise max_payload; do not remove required fields or replay the operation.

For S3 storage, set storage.backend=s3 and choose storage.provider=rustfs with rustfs.enabled=true for bundled storage, or storage.provider=external with rustfs.enabled=false, an S3 endpoint, and credentials for an existing service. Bucket bootstrap uses the RustFS rc client.

Coordinator and worker settings are separate:

  • Configure the control plane through coordinator.podSecurityContext and coordinator.containerSecurityContext.
  • Configure the execution plane through workerDefaults.podSecurityContext and workerDefaults.containerSecurityContext, with a named pool override only when its tool image requires one.

The coordinator defaults to UID/GID 10001, non-root execution, a read-only root filesystem, no Linux capabilities, and RuntimeDefault seccomp.

The worker uses a different boundary deliberately. Its supervisor starts as UID/GID 0, the root filesystem remains read-only, privilege escalation is disabled, and every capability is dropped except CHOWN, DAC_OVERRIDE, FOWNER, SETUID, and SETGID. The supervisor uses those capabilities to hand workspace ownership to an isolated IaC identity and to materialize or clean private run files. The chart sets that child identity through FORGEPLANE_WORKER_IAC_UID and FORGEPLANE_WORKER_IAC_GID in workerDefaults.extraEnv (both default to 10001). Do not replace this with a generic non-root context without replacing and testing the execution-identity mechanism.

ServiceAccount token automount is disabled. Workers additionally need chart-managed writable mounts for their temporary directory and workspace root. Keep those mounts and provider credentials out of the coordinator pod.

workerDefaults contains common worker runtime, probes, security, service, storage, and health settings. workerPools contains named fleets and per-pool overrides.

The default production worker image contains no IaC tools, so the default pool advertises an empty capability list. A tool-enabled pool must use an image that contains the named tools and must not advertise anything else.

workerDefaults:
grpc:
allowInsecure: false
runtime:
tempDir: /tmp/forgeplane/tmp
workspaceRoot: /var/lib/forgeplane/workspaces
workerPools:
default:
capabilities: []
tofu:
image:
repository: registry.example.com/forgeplane-worker-tofu
tag: "reviewed-release"
capabilities:
- tofu
extraEnv:
- name: FORGEPLANE_WORKER_IAC_UID
value: "10001"
- name: FORGEPLANE_WORKER_IAC_GID
value: "10001"
- name: FORGEPLANE_WORKER_POOL_NAME
value: tofu
autoscaling:
enabled: true

The chart fleet name does not set the worker’s registration pool. Set FORGEPLANE_WORKER_POOL_NAME to the pool required by templates. A pool-level extraEnv list replaces workerDefaults.extraEnv, so retain the IaC UID/GID entries.

Pools can independently select Deployment or StatefulSet, health Service exposure, resources, and horizontal scaling. Scaling is capacity-based Kubernetes scaling; the scheduler still requires a live, capability-matching worker and a valid assignment lease.

The repository’s OpenTofu 1.10.6 image is an opt-in Docker Compose validation image, not the production Helm worker. Build and publish a reviewed tool image, then set its exact repository and versioned tag in the pool values supplied with your release. The chart renders images as repository:tag and has no dedicated digest field. To pin a digest, keep repository as the image name and set tag to <version>@sha256:<digest>, producing name:version@sha256:<digest>. Apply the same form to coordinator, migrator, worker defaults, and any pool image overrides. Do not put an @sha256: digest in repository while the chart appends tag.

The chart’s secure gRPC path uses a coordinator server certificate, a worker client certificate, and a shared CA. With chart-managed certificates:

certManager:
enabled: true
coordinator:
grpc:
allowInsecure: false
requireClientCert: true
reflectionEnabled: false
workerDefaults:
grpc:
allowInsecure: false

When certificates are managed elsewhere, set the coordinator server Secret, worker client Secret, CA Secret, and worker server name explicitly. Enable insecure gRPC on both sides only for local plaintext testing.

Use S3-compatible storage for durable production artifacts and managed state. Coordinator /readyz checks PostgreSQL, NATS/JetStream, memory, and object storage; /healthz is process liveness only. Storage policy must permit the readiness probe’s create, multipart-write, and delete operations under its reserved prefix.

Forgeplane-managed state spans three coupled systems:

  1. PostgreSQL metadata and current-generation pointers;
  2. encrypted blobs in object storage; and
  3. every decrypt key ID referenced by retained ciphertext revisions.

Back up and restore these as one consistent, manifested recovery set. A database dump or bucket snapshot by itself is not recoverable state. Keep decrypt keys in separately protected escrow, define installation RPO/RTO targets, and perform restore drills before enabling mutation in a recovered environment.

The chart’s built-in egress NetworkPolicy selects the whole Forgeplane release: coordinator pods and every worker pool. networkPolicy.extraEgress therefore grants the same destinations to all of them. If that policy denies external egress, allow only destinations that every selected component may use. Use component- or pool-scoped external NetworkPolicies—or add and test chart support for separate selectors—when PostgreSQL and object storage must remain coordinator-only while workers receive DNS, coordinator gRPC, NATS, source, telemetry, or target access. Keep queue, storage, and SMTP credentials in Secrets.

When rotating an externally managed runtime Secret, update coordinator.existingSecretChecksum or use the supported reloader integration so coordinator and worker pods roll to the new material.

Before enabling production mutations, verify:

  • dependency ownership and backup responsibility are documented;
  • coordinator and worker images are reviewed and immutably pinned;
  • worker capabilities match the binaries inside each image;
  • gRPC certificates and client authentication are enabled;
  • readiness checks can perform their reserved storage operations; and
  • managed-state restore has been tested with the database, blobs, and decrypt keys together.