Kubernetes
The Helm chart is the supported way to run ScramDB on Kubernetes, for a single node and for a cluster. It is the same engine and the same image either way; the chart decides which shape you get and derives the values that would otherwise be yours to keep in step.
helm repo add scramdb https://charts.scramdb.com
helm repo update
helm install scramdb scramdb/scramdb
That is one Community node, no license, and it installs on any cluster including
kind and minikube. Everything below is how you turn it into what you actually
want.
Source and full values reference: github.com/scramdblabs/charts.
A three node cluster
Multi-node clustering is an Enterprise capability, so a cluster needs a license
before anything else. A node with a [cluster] section whose license resolves to
Community refuses to start, and the chart refuses to render that combination rather
than installing a crash loop.
kubectl create secret generic scramdb-license \
--from-literal=license-key="<your Enterprise or Trial token>"
helm install scramdb scramdb/scramdb \
--set architecture=cluster \
--set replicaCount=3 \
--set license.existingSecret=scramdb-license \
--set resourcesPreset=medium \
--set persistence.size=500Gi \
--set walArchive.destination=s3://your-bucket/scramdb-wal
Watch it form:
kubectl get pods -l app.kubernetes.io/instance=scramdb -w
A pod turns Ready once the cluster has formed. Then ask the cluster about itself, from any pod:
kubectl exec scramdb-0 -- psql -h 127.0.0.1 -U scramdb -d scramdb \
-c "SELECT node_id, role, state FROM scram.nodes ORDER BY node_id"
Expect one voter row per pod, each live; SHOW CLUSTER prints the same rows. The
formation lines (cluster transport: bound, swim: member up) log at info: install with
--set logLevel=info to see them in kubectl logs.
What the chart decides for you
Two values in a ScramDB cluster config have to agree with the topology, and both are the classic way a hand-written manifest goes stale. The chart computes them from the replica count instead of asking you to type them.
| Setting | Derived as | Why it is not yours to set |
|---|---|---|
bootstrap_expect | total voters minus one, summed across every region pool | It has to equal the number of other voters. A stale value does not misform silently: the node waits out group0_bootstrap_timeout and then refuses, naming what it waited for. |
columnar_replica | true at three voters or fewer, false above | At three nodes there is no smaller topology to dedicate an analytics tier out of, so every voter mirrors its own committed entries and answers analytics with no hop. Above that a learner tier buys real resource isolation instead. |
Both accept an explicit override (cluster.bootstrapExpect, cluster.columnarReplica)
for the case where you are joining a cluster the chart does not own.
Scaling up needs no config change at all: a node that finds an already-formed cluster
boots as an elastic joiner rather than founding a rival one, and the running cluster
adds it as a non-voting member and promotes it once it has caught up. bootstrap_expect
only ever decides who founds a cluster from nothing.
How a node finds its peers
There is no seed list. Every pod runs one identical config template, and the only per-pod substitution is its own name:
[cluster]
node_name = "${NODE_NAME}"
cluster_listen = "0.0.0.0:7190"
cluster_interactive_listen = "0.0.0.0:7191"
cluster_bulk_listen = "0.0.0.0:7192"
advertise_addr = "${NODE_NAME}.scramdb-headless.default.svc.cluster.local:7190"
bootstrap_expect = 2
discovery = ["dns", "swim"]
dns_name = "scramdb-headless.default.svc.cluster.local"
dns_refresh = "5s"
replication_factor = 3
An init container substitutes ${NODE_NAME} from the pod's metadata.name, read
through the downward API, and writes the rendered file to a shared volume. The serving
container then runs scramdb -c against it, with no shell of its own.
The headless Service is what makes that work, and it sets
publishNotReadyAddresses: true deliberately. A pod is only Ready once it serves
queries, which in cluster mode only happens after the cluster has formed. If DNS
published Ready addresses only, no pod on a fresh cluster could ever see another, and
the cluster would never form at all.
Sizing
resourcesPreset picks a shape and resources overrides it outright.
| Preset | CPU | Memory |
|---|---|---|
dev (default) | 1 | 2Gi |
small | 4 | 8Gi |
medium | 16 | 32Gi |
large | 32 | 64Gi |
xlarge | 64 | 128Gi |
medium is the reference production shape: 16 vCPU and 32 GiB, matching the
c6a.4xlarge instance ScramDB's published benchmarks are measured on. The default is
dev so a first install schedules anywhere; move to medium or above for a real
workload.
Every preset sets requests equal to limits, which gives the pod Guaranteed QoS. That
is what lets a kubelet running the static CPU manager policy pin exclusive cores per
node, the best of the options in Parallelism. Under the
default CPU manager policy the engine still resolves the right worker count from the
CFS quota, so an integer limit is never wrong, only less predictable at the tail.
Memory needs no arithmetic on your side. buffer_pool_percent and
execution_memory_percent stay at 0, which means the engine sizes both from the
container's own cgroup memory limit. The same values are correct at 2Gi and at 512Gi.
Storage
persistence:
enabled: true
storageClass: "" # the cluster's default class
accessModes: [ReadWriteOnce]
size: 100Gi
An empty storageClass uses whatever the cluster's default is, which is what lets one
values file work on EKS, GKE, AKS and on-prem alike. NVMe backed classes are strongly
recommended: the default I/O backend uses direct I/O, bypassing the OS page cache, so
scans reward sequential throughput.
Everything the engine writes lives under this one volume: the data, the WAL, the spill
directory, the compiled-artifact cache and the UDF cache all derive from
storage.basedir. That is also what lets the container run with a read-only root
filesystem.
WAL archiving
WAL archiving is on by default, which is what makes point-in-time recovery and branching work. In a cluster, point it at storage every node can reach:
walArchive:
destination: s3://your-bucket/scramdb-wal # or gs:// or az:// or a shared file://
Only the current leader ships segments, and a failed-over leader resumes from the destination's own listing. With a node-local destination each new leader starts a fresh archive, so no node ever holds a complete history. The chart warns about this at install time rather than letting it look like it worked.
Connecting
Every node is an equal peer serving the same consistent database, so any pod is a correct endpoint and the client Service load balances across all of them.
postgresql://scramdb@scramdb.default.svc.cluster.local:5432/scramdb
The chart turns authentication on and seeds the bootstrap superuser's password at first boot. Read it back with:
kubectl get secret scramdb-auth -o jsonpath="{.data.password}" | base64 -d
If you do not supply auth.password or auth.existingSecret, the chart generates one
and keeps it across upgrades. The seeded password applies only while the superuser has
no credential at all, so redeploying a pod or rotating the Secret can never reset a
password you set with ALTER ROLE, and can never lock anyone out.
auth.seedMethod picks how it reaches the engine. file (the default) mounts the
Secret and reads it once, so the value never appears in the pod spec or
/proc/<pid>/environ. env passes the literal value, which is visible in
kubectl describe pod. random lets the engine generate one and log it exactly once at
first boot. Exactly one may be set; the engine refuses to start on more than one, and
the chart only ever emits one.
Ports
Every port a node opens is published on the headless Service, so any of them can be
reached per-pod. service.* decides which the client Service carries.
| Port | What listens | On the client Service |
|---|---|---|
| 5432 | The PostgreSQL wire protocol | Always |
| 7190 | The cluster transport's control traffic | Never, it is peer to peer traffic |
| 7191 | The cluster transport's interactive traffic | Never, it is peer to peer traffic |
| 7192 | The cluster transport's bulk traffic | Never, it is peer to peer traffic |
| 9090 | /metrics and /health | service.exposeMetrics |
| 9191 | The Semantic AI MCP server | service.exposeMcp, on by default |
With networkPolicy.enabled, the three transport ports are let through between the chart's
own pods only, in both directions, whatever networkPolicy.allowExternal says. The chart sets
all three transport addresses in the node's config, so a taken port stops the pod instead of
moving a lane to a port no Service or policy carries. Ports has the
firewall rules for everything outside the cluster.
The MCP server is started in-process by the engine from the packages baked into the
image, so every node opens 9191 whether or not you publish it. The chart sets
MCP_AUTH=basic, which maps each request's HTTP Basic credentials to a real database
role, so the tool surface is exactly as protected as pgwire. The alternative,
mcp.auth=env, leaves it with no HTTP authentication at all; the chart refuses to
render that combination behind a LoadBalancer or a NodePort.
Multi-region
One StatefulSet per region, all sharing one headless Service, forming one flat cluster.
architecture: cluster
regions:
- name: eu
zone: eu-central-1a
replicas: 3
nodeSelector:
topology.kubernetes.io/region: eu-central-1
- name: us
zone: us-east-1a
replicas: 3
nodeSelector:
topology.kubernetes.io/region: us-east-1
A table created through a node labelled eu keeps its voters in eu and commits at
region-local quorum latency; ALTER TABLE t SET (home_region = 'us') re-homes it live.
bootstrap_expect is derived from the sum of the pools. See
Multi-region for what the labels buy.
An empty or unset label is unset, never a region literally named "", so a pool with no
name-derived region behaves exactly like an unlabelled deployment.
Learners
A learner is a permanently non-voting member: it hosts replicated data for local analytical reads, owns no write-serving buckets, and is never promoted to a voter.
learners:
enabled: true
replicas: 2
This is the right shape at five nodes or more, where a dedicated tier buys hard resource
isolation between the transactional and analytical workloads. Below that, leave
cluster.columnarReplica on its derived true and let every voter answer analytics
locally instead. The two are mutually exclusive per node, and the engine refuses the
combination at boot.
Scaling
helm upgrade scramdb scramdb/scramdb --reuse-values --set replicaCount=5
Adding a node needs no command against the database. Removing one is safe too: a terminating pod hands its bucket ownership and its voter seat back before it exits.
terminationGracePeriodSeconds defaults to 120 for exactly that reason. The membership
hand-back, the query drain and the JIT quiesce each get their own 10 second budget on top
of the preStop sleep, so a worst case graceful stop wants roughly forty seconds. A pod
killed before it finishes stays a configured voter the cluster keeps counting, which is
what costs quorum on the next scale-down.
To drain a node explicitly first, from any pod:
ALTER CLUSTER DRAIN 'scramdb-2';
See Scaling for what that statement does and does not fan out.
Changing the replica count re-derives bootstrap_expect, so the pods roll once. That is
real reconfiguration work rather than a no-op, and it is the price of a value that cannot
go stale.
Security
The container security context satisfies the restricted Pod Security Standard: non-root
at UID 1001, no privilege escalation, all capabilities dropped, a read-only root
filesystem and the RuntimeDefault seccomp profile. fsGroup is what makes the data
volume writable, which works whatever UID the process ends up as.
The UID is a pinned contract in the image. An older image declares its user by name
instead, which a kubelet cannot verify, so a pod asking for runAsNonRoot: true is
rejected outright. Running one of those needs
--set containerSecurityContext.runAsNonRoot=false with runAsUser and runAsGroup
cleared.
Other levers:
tls.enabledwithtls.existingSecretturns on server-side TLS for pgwire. A missing or mismatched certificate fails startup rather than falling back to plaintext. See TLS.cluster.tls.enabledadds mutual TLS between nodes on the cluster transport, from the certificate incluster.tls.existingSecret. That one shared certificate must carry every node name (the pod names) and the pods' stable DNS names, never their IPs: the chart's install notes (helm status) print the exact list and the commands that make the certificate and the secret. See TLS between cluster nodes.hba.rulestakes apg_hba.confstyle rule file. Empty uses the built-in default: trust from localhost, password everywhere else. See Authentication.networkPolicy.enabledrestricts the client and MCP ports to selectors you name. The cluster transport's three ports (7190,7191,7192) are always restricted to the chart's own pods regardless.auth.enabled=falseruns the node with--pg-no-auth, trusting every connection with no password check. Only for a network you already control.
The cost of the hardened defaults
The RuntimeDefault seccomp profile blocks the io_uring syscalls. ScramDB detects this
at startup, falls back to the epoll I/O path and says so:
morsel workers: io_uring unavailable (Operation not permitted (os error 1));
falling back to epoll
That line is expected on a hardened pod, not a misconfiguration, and the engine keeps
working. What it costs you is the io_uring tier of the I/O ladder, which matters for
scan-heavy and bulk-load work. The two are mutually exclusive as things stand: a pod that
satisfies the restricted Pod Security Standard cannot use io_uring, and a pod that uses
io_uring needs a seccomp profile permitting those syscalls, which is a deliberate
exception to make with your platform team rather than a chart default.
Dropping all capabilities has a second, smaller effect. Background pools ask for a lower
CPU priority and cannot get it without CAP_SYS_NICE, so they run at the default priority
and log one line saying so. Query correctness is unaffected; the separation between
background maintenance and query workers is simply weaker than configured.
priority probe: pool 'core-pinned-workers' requested nice -20 but achieved 0 -
elevation degraded (missing CAP_SYS_NICE? default in containers)
Observability
metrics:
serviceMonitor:
enabled: true
prometheusRule:
enabled: true
The ServiceMonitor scrapes through the headless Service, so every node is its own Prometheus target, which is the only way per-node cluster counters mean anything. The PrometheusRule covers the failures that have a counter: a quarantined shard group, WAL archiving falling behind, peers flapping, an expiring license, execution-memory stalls. An absent metric on this endpoint means not measured, never zero, so nothing alerts on absence. See Observability for the full metric reference.
Both need the Prometheus Operator CRDs installed in the cluster.
Probes
Readiness and liveness are both pg_isready, which proves the node accepts client
connections. In cluster mode that only becomes true after the cluster has formed, which
is what the startup probe is for: it tolerates ten minutes by default, comfortably past
group0_bootstrap_timeout, so a node waiting for its peers is never mistaken for a node
that failed.
Troubleshooting
Pods stay at CreateContainerConfigError. The license Secret is missing.
kubectl describe pod names exactly which Secret it wanted. This is the kubelet catching
the same refusal a Community-licensed process would log on its own stderr, one layer up.
Pods restart before they are ever Ready. Check the startup probe budget against
cluster.group0BootstrapTimeout. A cluster that cannot reach its expected peer count
waits out that timeout and then refuses loudly, naming what it was waiting for; the log
line is the answer, not the probe.
A node refuses to start naming bootstrap_expect. Its expected peer count exceeds
what discovery can supply. Check that replicaCount matches reality and that you have
not set cluster.seeds alongside a derived bootstrap_expect.
The cluster formed but one pod is alone. Look for swim: member up lines. A pod that
never sees its peers usually cannot resolve the headless Service; confirm
clusterDomain matches your cluster's actual DNS suffix.
Writing your own manifests
The engine repository carries the hand-written equivalents under k8s/, as a kind
development loop and as the blueprint the chart was written from. They are not the
supported deployment path: the chart exists because bootstrap_expect, the region pools
and the shutdown budget are all things a hand-written manifest gets wrong by drifting,
not by being written wrong once.
Next
- Distributed cluster for what a cluster guarantees.
- Connecting to a cluster for the client side.
- Scaling for adding and removing nodes.
- Requirements for sizing beyond the presets.