Skip to main content

Kubernetes

The Helm chart is the supported way to run ScramDB on Kubernetes, for a single node and for a cluster. It is the same engine and the same image either way; the chart decides which shape you get and derives the values that would otherwise be yours to keep in step.

helm repo add scramdb https://charts.scramdb.com
helm repo update
helm install scramdb scramdb/scramdb

That is one Community node, no license, and it installs on any cluster including kind and minikube. Everything below is how you turn it into what you actually want.

Source and full values reference: github.com/scramdblabs/charts.

A three node cluster​

Multi-node clustering is an Enterprise capability, so a cluster needs a license before anything else. A node with a [cluster] section whose license resolves to Community refuses to start, and the chart refuses to render that combination rather than installing a crash loop.

kubectl create secret generic scramdb-license \
--from-literal=license-key="<your Enterprise or Trial token>"

helm install scramdb scramdb/scramdb \
--set architecture=cluster \
--set replicaCount=3 \
--set license.existingSecret=scramdb-license \
--set resourcesPreset=medium \
--set persistence.size=500Gi \
--set walArchive.destination=s3://your-bucket/scramdb-wal

Watch it form:

kubectl get pods -l app.kubernetes.io/instance=scramdb -w

A pod turns Ready once the cluster has formed. Then ask the cluster about itself, from any pod:

kubectl exec scramdb-0 -- psql -h 127.0.0.1 -U scramdb -d scramdb \
-c "SELECT node_id, role, state FROM scram.nodes ORDER BY node_id"

Expect one voter row per pod, each live; SHOW CLUSTER prints the same rows. The formation lines (cluster transport: bound, swim: member up) log at info: install with --set logLevel=info to see them in kubectl logs.

What the chart decides for you​

Two values in a ScramDB cluster config have to agree with the topology, and both are the classic way a hand-written manifest goes stale. The chart computes them from the replica count instead of asking you to type them.

SettingDerived asWhy it is not yours to set
bootstrap_expecttotal voters minus one, summed across every region poolIt has to equal the number of other voters. A stale value does not misform silently: the node waits out group0_bootstrap_timeout and then refuses, naming what it waited for.
columnar_replicatrue at three voters or fewer, false aboveAt three nodes there is no smaller topology to dedicate an analytics tier out of, so every voter mirrors its own committed entries and answers analytics with no hop. Above that a learner tier buys real resource isolation instead.

Both accept an explicit override (cluster.bootstrapExpect, cluster.columnarReplica) for the case where you are joining a cluster the chart does not own.

Scaling up needs no config change at all: a node that finds an already-formed cluster boots as an elastic joiner rather than founding a rival one, and the running cluster adds it as a non-voting member and promotes it once it has caught up. bootstrap_expect only ever decides who founds a cluster from nothing.

How a node finds its peers​

There is no seed list. Every pod runs one identical config template, and the only per-pod substitution is its own name:

[cluster]
node_name = "${NODE_NAME}"
cluster_listen = "0.0.0.0:7190"
cluster_interactive_listen = "0.0.0.0:7191"
cluster_bulk_listen = "0.0.0.0:7192"
advertise_addr = "${NODE_NAME}.scramdb-headless.default.svc.cluster.local:7190"
bootstrap_expect = 2
discovery = ["dns", "swim"]
dns_name = "scramdb-headless.default.svc.cluster.local"
dns_refresh = "5s"
replication_factor = 3

An init container substitutes ${NODE_NAME} from the pod's metadata.name, read through the downward API, and writes the rendered file to a shared volume. The serving container then runs scramdb -c against it, with no shell of its own.

The headless Service is what makes that work, and it sets publishNotReadyAddresses: true deliberately. A pod is only Ready once it serves queries, which in cluster mode only happens after the cluster has formed. If DNS published Ready addresses only, no pod on a fresh cluster could ever see another, and the cluster would never form at all.

Sizing​

resourcesPreset picks a shape and resources overrides it outright.

PresetCPUMemory
dev (default)12Gi
small48Gi
medium1632Gi
large3264Gi
xlarge64128Gi

medium is the reference production shape: 16 vCPU and 32 GiB, matching the c6a.4xlarge instance ScramDB's published benchmarks are measured on. The default is dev so a first install schedules anywhere; move to medium or above for a real workload.

Every preset sets requests equal to limits, which gives the pod Guaranteed QoS. That is what lets a kubelet running the static CPU manager policy pin exclusive cores per node, the best of the options in Parallelism. Under the default CPU manager policy the engine still resolves the right worker count from the CFS quota, so an integer limit is never wrong, only less predictable at the tail.

Memory needs no arithmetic on your side. buffer_pool_percent and execution_memory_percent stay at 0, which means the engine sizes both from the container's own cgroup memory limit. The same values are correct at 2Gi and at 512Gi.

Storage​

persistence:
enabled: true
storageClass: "" # the cluster's default class
accessModes: [ReadWriteOnce]
size: 100Gi

An empty storageClass uses whatever the cluster's default is, which is what lets one values file work on EKS, GKE, AKS and on-prem alike. NVMe backed classes are strongly recommended: the default I/O backend uses direct I/O, bypassing the OS page cache, so scans reward sequential throughput.

Everything the engine writes lives under this one volume: the data, the WAL, the spill directory, the compiled-artifact cache and the UDF cache all derive from storage.basedir. That is also what lets the container run with a read-only root filesystem.

WAL archiving​

WAL archiving is on by default, which is what makes point-in-time recovery and branching work. In a cluster, point it at storage every node can reach:

walArchive:
destination: s3://your-bucket/scramdb-wal # or gs:// or az:// or a shared file://

Only the current leader ships segments, and a failed-over leader resumes from the destination's own listing. With a node-local destination each new leader starts a fresh archive, so no node ever holds a complete history. The chart warns about this at install time rather than letting it look like it worked.

Connecting​

Every node is an equal peer serving the same consistent database, so any pod is a correct endpoint and the client Service load balances across all of them.

postgresql://scramdb@scramdb.default.svc.cluster.local:5432/scramdb

The chart turns authentication on and seeds the bootstrap superuser's password at first boot. Read it back with:

kubectl get secret scramdb-auth -o jsonpath="{.data.password}" | base64 -d

If you do not supply auth.password or auth.existingSecret, the chart generates one and keeps it across upgrades. The seeded password applies only while the superuser has no credential at all, so redeploying a pod or rotating the Secret can never reset a password you set with ALTER ROLE, and can never lock anyone out.

auth.seedMethod picks how it reaches the engine. file (the default) mounts the Secret and reads it once, so the value never appears in the pod spec or /proc/<pid>/environ. env passes the literal value, which is visible in kubectl describe pod. random lets the engine generate one and log it exactly once at first boot. Exactly one may be set; the engine refuses to start on more than one, and the chart only ever emits one.

Ports​

Every port a node opens is published on the headless Service, so any of them can be reached per-pod. service.* decides which the client Service carries.

PortWhat listensOn the client Service
5432The PostgreSQL wire protocolAlways
7190The cluster transport's control trafficNever, it is peer to peer traffic
7191The cluster transport's interactive trafficNever, it is peer to peer traffic
7192The cluster transport's bulk trafficNever, it is peer to peer traffic
9090/metrics and /healthservice.exposeMetrics
9191The Semantic AI MCP serverservice.exposeMcp, on by default

With networkPolicy.enabled, the three transport ports are let through between the chart's own pods only, in both directions, whatever networkPolicy.allowExternal says. The chart sets all three transport addresses in the node's config, so a taken port stops the pod instead of moving a lane to a port no Service or policy carries. Ports has the firewall rules for everything outside the cluster.

The MCP server is started in-process by the engine from the packages baked into the image, so every node opens 9191 whether or not you publish it. The chart sets MCP_AUTH=basic, which maps each request's HTTP Basic credentials to a real database role, so the tool surface is exactly as protected as pgwire. The alternative, mcp.auth=env, leaves it with no HTTP authentication at all; the chart refuses to render that combination behind a LoadBalancer or a NodePort.

Multi-region​

One StatefulSet per region, all sharing one headless Service, forming one flat cluster.

architecture: cluster
regions:
- name: eu
zone: eu-central-1a
replicas: 3
nodeSelector:
topology.kubernetes.io/region: eu-central-1
- name: us
zone: us-east-1a
replicas: 3
nodeSelector:
topology.kubernetes.io/region: us-east-1

A table created through a node labelled eu keeps its voters in eu and commits at region-local quorum latency; ALTER TABLE t SET (home_region = 'us') re-homes it live. bootstrap_expect is derived from the sum of the pools. See Multi-region for what the labels buy.

An empty or unset label is unset, never a region literally named "", so a pool with no name-derived region behaves exactly like an unlabelled deployment.

Learners​

A learner is a permanently non-voting member: it hosts replicated data for local analytical reads, owns no write-serving buckets, and is never promoted to a voter.

learners:
enabled: true
replicas: 2

This is the right shape at five nodes or more, where a dedicated tier buys hard resource isolation between the transactional and analytical workloads. Below that, leave cluster.columnarReplica on its derived true and let every voter answer analytics locally instead. The two are mutually exclusive per node, and the engine refuses the combination at boot.

Scaling​

helm upgrade scramdb scramdb/scramdb --reuse-values --set replicaCount=5

Adding a node needs no command against the database. Removing one is safe too: a terminating pod hands its bucket ownership and its voter seat back before it exits.

terminationGracePeriodSeconds defaults to 120 for exactly that reason. The membership hand-back, the query drain and the JIT quiesce each get their own 10 second budget on top of the preStop sleep, so a worst case graceful stop wants roughly forty seconds. A pod killed before it finishes stays a configured voter the cluster keeps counting, which is what costs quorum on the next scale-down.

To drain a node explicitly first, from any pod:

ALTER CLUSTER DRAIN 'scramdb-2';

See Scaling for what that statement does and does not fan out.

Changing the replica count re-derives bootstrap_expect, so the pods roll once. That is real reconfiguration work rather than a no-op, and it is the price of a value that cannot go stale.

Security​

The container security context satisfies the restricted Pod Security Standard: non-root at UID 1001, no privilege escalation, all capabilities dropped, a read-only root filesystem and the RuntimeDefault seccomp profile. fsGroup is what makes the data volume writable, which works whatever UID the process ends up as.

Images built before the UID was pinned

The UID is a pinned contract in the image. An older image declares its user by name instead, which a kubelet cannot verify, so a pod asking for runAsNonRoot: true is rejected outright. Running one of those needs --set containerSecurityContext.runAsNonRoot=false with runAsUser and runAsGroup cleared.

Other levers:

  • tls.enabled with tls.existingSecret turns on server-side TLS for pgwire. A missing or mismatched certificate fails startup rather than falling back to plaintext. See TLS.
  • cluster.tls.enabled adds mutual TLS between nodes on the cluster transport, from the certificate in cluster.tls.existingSecret. That one shared certificate must carry every node name (the pod names) and the pods' stable DNS names, never their IPs: the chart's install notes (helm status) print the exact list and the commands that make the certificate and the secret. See TLS between cluster nodes.
  • hba.rules takes a pg_hba.conf style rule file. Empty uses the built-in default: trust from localhost, password everywhere else. See Authentication.
  • networkPolicy.enabled restricts the client and MCP ports to selectors you name. The cluster transport's three ports (7190, 7191, 7192) are always restricted to the chart's own pods regardless.
  • auth.enabled=false runs the node with --pg-no-auth, trusting every connection with no password check. Only for a network you already control.

The cost of the hardened defaults​

The RuntimeDefault seccomp profile blocks the io_uring syscalls. ScramDB detects this at startup, falls back to the epoll I/O path and says so:

morsel workers: io_uring unavailable (Operation not permitted (os error 1));
falling back to epoll

That line is expected on a hardened pod, not a misconfiguration, and the engine keeps working. What it costs you is the io_uring tier of the I/O ladder, which matters for scan-heavy and bulk-load work. The two are mutually exclusive as things stand: a pod that satisfies the restricted Pod Security Standard cannot use io_uring, and a pod that uses io_uring needs a seccomp profile permitting those syscalls, which is a deliberate exception to make with your platform team rather than a chart default.

Dropping all capabilities has a second, smaller effect. Background pools ask for a lower CPU priority and cannot get it without CAP_SYS_NICE, so they run at the default priority and log one line saying so. Query correctness is unaffected; the separation between background maintenance and query workers is simply weaker than configured.

priority probe: pool 'core-pinned-workers' requested nice -20 but achieved 0 -
elevation degraded (missing CAP_SYS_NICE? default in containers)

Observability​

metrics:
serviceMonitor:
enabled: true
prometheusRule:
enabled: true

The ServiceMonitor scrapes through the headless Service, so every node is its own Prometheus target, which is the only way per-node cluster counters mean anything. The PrometheusRule covers the failures that have a counter: a quarantined shard group, WAL archiving falling behind, peers flapping, an expiring license, execution-memory stalls. An absent metric on this endpoint means not measured, never zero, so nothing alerts on absence. See Observability for the full metric reference.

Both need the Prometheus Operator CRDs installed in the cluster.

Probes​

Readiness and liveness are both pg_isready, which proves the node accepts client connections. In cluster mode that only becomes true after the cluster has formed, which is what the startup probe is for: it tolerates ten minutes by default, comfortably past group0_bootstrap_timeout, so a node waiting for its peers is never mistaken for a node that failed.

Troubleshooting​

Pods stay at CreateContainerConfigError. The license Secret is missing. kubectl describe pod names exactly which Secret it wanted. This is the kubelet catching the same refusal a Community-licensed process would log on its own stderr, one layer up.

Pods restart before they are ever Ready. Check the startup probe budget against cluster.group0BootstrapTimeout. A cluster that cannot reach its expected peer count waits out that timeout and then refuses loudly, naming what it was waiting for; the log line is the answer, not the probe.

A node refuses to start naming bootstrap_expect. Its expected peer count exceeds what discovery can supply. Check that replicaCount matches reality and that you have not set cluster.seeds alongside a derived bootstrap_expect.

The cluster formed but one pod is alone. Look for swim: member up lines. A pod that never sees its peers usually cannot resolve the headless Service; confirm clusterDomain matches your cluster's actual DNS suffix.

Writing your own manifests​

The engine repository carries the hand-written equivalents under k8s/, as a kind development loop and as the blueprint the chart was written from. They are not the supported deployment path: the chart exists because bootstrap_expect, the region pools and the shutdown budget are all things a hand-written manifest gets wrong by drifting, not by being written wrong once.

Next​