Forgeplane Helm deployment on Kubernetes
The Forgeplane chart deploys the coordinator, database initialization, and one or more named worker pools. It can install development-sized dependencies, or connect to services you operate separately. Production deployments must make explicit choices for durability, TLS, artifact storage, managed state, and backups.
Prepare the installation
Section titled “Prepare the installation”You need:
- Kubernetes and Helm 3;
- signed coordinator, worker, and migrator images for the release;
- PostgreSQL, NATS JetStream, and local or S3-compatible artifact storage;
- cert-manager CRDs when
certManager.enabled=true; and - a runtime Secret containing the keys required by enabled database, queue, storage, bootstrap, mail, and metrics modes.
Build chart dependencies before installing from a source checkout:
make chart-dependencieshelm install forgeplane charts/forgeplane \ --namespace forgeplane \ --create-namespace \ -f values-prod.yamlProduction secret defaults are intentionally empty. A raw render without suitable values should fail rather than emit placeholder credentials.
Pre-production builds without releases
Section titled “Pre-production builds without releases”Maintainers can explicitly publish a development build from main without a Git tag or GitHub release. Each workflow attempt gets a unique 0.0.0-dev.<run-id>.<attempt> version and a successful build receipt recording the source commit, image digests, and chart version/OCI digest. The publisher does not update latest or deploy anything.
Follow the repository’s development-build runbook for publication approval, signature verification, private registry access, and reviewed GitOps pins. Use one complete successful receipt, not artifacts from failed or mixed attempts; deployment approval remains separate.
Choose dependency ownership
Section titled “Choose dependency ownership”The chart can install these dependencies or connect to external services:
| Service | Subchart | Enable value | Production choice |
|---|---|---|---|
| PostgreSQL | bitnami/postgresql |
postgresql.enabled |
Prefer a managed or independently operated database with tested backups. |
| NATS JetStream | nats/nats |
nats.enabled |
Configure authentication, TLS, persistence, and clustering deliberately. |
| RustFS | rustfs/rustfs |
rustfs.enabled |
S3-compatible option for artifact and state blobs. |
coordinator.existingSecret replaces only the Forgeplane runtime Secret. It does not supply credentials required by enabled PostgreSQL, NATS, or RustFS subcharts.
For external PostgreSQL, use postgresql.connection.sslMode=verify-full and a hostname that matches the server certificate. The coordinator and migrator use the image trust store; the chart does not add a private database CA or client certificate automatically.
Every external NATS server must use max_payload >= 4194304 so the existing 2
MiB structured-output contract and its signed execution-outcome envelope fit. If
Forgeplane reports the observed limit, raise max_payload; do not remove required
fields or replay the operation.
Choose object storage
Section titled “Choose object storage”For S3 storage, set storage.backend=s3 and choose storage.provider=rustfs with rustfs.enabled=true for bundled storage, or storage.provider=external with rustfs.enabled=false, an S3 endpoint, and credentials for an existing service. Bucket bootstrap uses the RustFS rc client.
Configure coordinator and worker security
Section titled “Configure coordinator and worker security”Coordinator and worker settings are separate:
- Configure the control plane through
coordinator.podSecurityContextandcoordinator.containerSecurityContext. - Configure the execution plane through
workerDefaults.podSecurityContextandworkerDefaults.containerSecurityContext, with a named pool override only when its tool image requires one.
The coordinator defaults to UID/GID 10001, non-root execution, a read-only root filesystem, no Linux capabilities, and RuntimeDefault seccomp.
The worker uses a different boundary deliberately. Its supervisor starts as UID/GID 0, the root filesystem remains read-only, privilege escalation is disabled, and every capability is dropped except CHOWN, DAC_OVERRIDE, FOWNER, SETUID, and SETGID. The supervisor uses those capabilities to hand workspace ownership to an isolated IaC identity and to materialize or clean private run files. The chart sets that child identity through FORGEPLANE_WORKER_IAC_UID and FORGEPLANE_WORKER_IAC_GID in workerDefaults.extraEnv (both default to 10001). Do not replace this with a generic non-root context without replacing and testing the execution-identity mechanism.
ServiceAccount token automount is disabled. Workers additionally need chart-managed writable mounts for their temporary directory and workspace root. Keep those mounts and provider credentials out of the coordinator pod.
Define capability-specific worker pools
Section titled “Define capability-specific worker pools”workerDefaults contains common worker runtime, probes, security, service, storage, and health settings. workerPools contains named fleets and per-pool overrides.
The default production worker image contains no IaC tools, so the default pool advertises an empty capability list. A tool-enabled pool must use an image that contains the named tools and must not advertise anything else.
workerDefaults: grpc: allowInsecure: false runtime: tempDir: /tmp/forgeplane/tmp workspaceRoot: /var/lib/forgeplane/workspaces
workerPools: default: capabilities: []
tofu: image: repository: registry.example.com/forgeplane-worker-tofu tag: "reviewed-release" capabilities: - tofu extraEnv: - name: FORGEPLANE_WORKER_IAC_UID value: "10001" - name: FORGEPLANE_WORKER_IAC_GID value: "10001" - name: FORGEPLANE_WORKER_POOL_NAME value: tofu autoscaling: enabled: trueThe chart fleet name does not set the worker’s registration pool. Set FORGEPLANE_WORKER_POOL_NAME to the pool required by templates. A pool-level extraEnv list replaces workerDefaults.extraEnv, so retain the IaC UID/GID entries.
Pools can independently select Deployment or StatefulSet, health Service exposure, resources, and horizontal scaling. Scaling is capacity-based Kubernetes scaling; the scheduler still requires a live, capability-matching worker and a valid assignment lease.
The repository’s OpenTofu 1.10.6 image is an opt-in Docker Compose validation image, not the production Helm worker. Build and publish a reviewed tool image, then set its exact repository and versioned tag in the pool values supplied with your release. The chart renders images as repository:tag and has no dedicated digest field. To pin a digest, keep repository as the image name and set tag to <version>@sha256:<digest>, producing name:version@sha256:<digest>. Apply the same form to coordinator, migrator, worker defaults, and any pool image overrides. Do not put an @sha256: digest in repository while the chart appends tag.
Configure gRPC transport security
Section titled “Configure gRPC transport security”The chart’s secure gRPC path uses a coordinator server certificate, a worker client certificate, and a shared CA. With chart-managed certificates:
certManager: enabled: true
coordinator: grpc: allowInsecure: false requireClientCert: true reflectionEnabled: false
workerDefaults: grpc: allowInsecure: falseWhen certificates are managed elsewhere, set the coordinator server Secret, worker client Secret, CA Secret, and worker server name explicitly. Enable insecure gRPC on both sides only for local plaintext testing.
Plan readiness and recovery
Section titled “Plan readiness and recovery”Use S3-compatible storage for durable production artifacts and managed state. Coordinator /readyz checks PostgreSQL, NATS/JetStream, memory, and object storage; /healthz is process liveness only. Storage policy must permit the readiness probe’s create, multipart-write, and delete operations under its reserved prefix.
Forgeplane-managed state spans three coupled systems:
- PostgreSQL metadata and current-generation pointers;
- encrypted blobs in object storage; and
- every decrypt key ID referenced by retained ciphertext revisions.
Back up and restore these as one consistent, manifested recovery set. A database dump or bucket snapshot by itself is not recoverable state. Keep decrypt keys in separately protected escrow, define installation RPO/RTO targets, and perform restore drills before enabling mutation in a recovered environment.
Configure network and secret operations
Section titled “Configure network and secret operations”The chart’s built-in egress NetworkPolicy selects the whole Forgeplane release: coordinator pods and every worker pool. networkPolicy.extraEgress therefore grants the same destinations to all of them. If that policy denies external egress, allow only destinations that every selected component may use. Use component- or pool-scoped external NetworkPolicies—or add and test chart support for separate selectors—when PostgreSQL and object storage must remain coordinator-only while workers receive DNS, coordinator gRPC, NATS, source, telemetry, or target access. Keep queue, storage, and SMTP credentials in Secrets.
When rotating an externally managed runtime Secret, update coordinator.existingSecretChecksum or use the supported reloader integration so coordinator and worker pods roll to the new material.
Production checklist
Section titled “Production checklist”Before enabling production mutations, verify:
- dependency ownership and backup responsibility are documented;
- coordinator and worker images are reviewed and immutably pinned;
- worker capabilities match the binaries inside each image;
- gRPC certificates and client authentication are enabled;
- readiness checks can perform their reserved storage operations; and
- managed-state restore has been tested with the database, blobs, and decrypt keys together.