Skip to main content

Observability

By the end of this page you will know how to scrape ScramDB's metrics, check liveness, and configure logging, with the exact set of metric names that exist today so you do not build a dashboard around a counter that isn't there.

The metrics and health endpoint​

ScramDB serves both /metrics and /health on the same HTTP listener, one port, one process.

SettingDefaultHow to set it
Port9090[general] metrics_port in TOML, or --metrics-port <PORT> on the command line. Set to 0 to disable the endpoint entirely.
Bind address0.0.0.0 (all interfaces)Not configurable. If you run ScramDB without a firewall or network policy in front of it, this port is reachable from outside the host.

/metrics​

curl -s http://127.0.0.1:9090/metrics | head -20

Expected output: Prometheus text exposition format (Content-Type: text/plain; version=0.0.4; charset=utf-8), starting with lines like # HELP scramdb_queries_total .... If it fails: connection refused means either the server isn't running, metrics_port is set to 0, or a firewall is blocking the port. Remember the listener binds 0.0.0.0, so "connection refused from outside the host" is more likely a firewall rule than the endpoint being loopback-only.

Any path other than /metrics or /health returns 404 Not Found.

/health​

curl -s http://127.0.0.1:9090/health

Expected output: ok (plain text, no JSON, no fields, HTTP 200). On a cluster node whose commit protocol has stopped after a failure (scramdb_cluster_dilith_host_stopped at 1), the answer is HTTP 503 Service Unavailable with the body the commit protocol stopped on this node: that node can no longer commit transactions and must be restarted, so a load balancer or orchestrator probing /health takes it out of service. There is no separate liveness/readiness distinction. Beyond that one case it does not check database or storage health internally: a 200 ok proves the process and its HTTP listener are up, not that queries are succeeding. For query success/failure, use the scramdb_queries_ok / scramdb_queries_err counters on /metrics instead.

The endpoint reports on itself on the same /metrics body, always emitted:

MetricTypeMeaning
scramdb_metrics_endpoint_requests_totalcounterRequests the metrics endpoint answered.
scramdb_metrics_endpoint_request_timeouts_totalcounterClients disconnected because their request did not arrive in time.
scramdb_metrics_endpoint_response_timeouts_totalcounterClients disconnected because they did not take their response in time.
scramdb_metrics_endpoint_client_errors_totalcounterClient connections that failed before their answer was written.
scramdb_metrics_endpoint_accept_failures_totalcounterConnections the metrics endpoint could not accept.
scramdb_metrics_endpoint_render_secondshistogramTime to render one /metrics answer.

Metric reference​

Every metric named on this page is really emitted on /metrics, so you can build a dashboard from these names without checking first. The reverse does not hold: the endpoint also carries deep diagnostic counters that are not part of the supported surface and can change between releases. The families below are the ones worth alerting on and graphing.

Core query and connection metrics (always present)​

MetricTypeMeaning
scramdb_queries_totalcounterTotal queries received.
scramdb_queries_okcounterQueries that completed successfully.
scramdb_queries_errcounterQueries that errored.
scramdb_queries_activegaugeQueries currently executing.
scramdb_query_avg_usgaugeRolling average query latency in microseconds.
scramdb_connections_totalcounterTotal connections accepted since startup.
scramdb_connections_activegaugeCurrently open connections.
scramdb_bytes_sentcounterTotal bytes sent to clients.
scramdb_bytes_receivedcounterTotal bytes received from clients.
scramdb_uptime_secondsgaugeSeconds since the process started.

Execution-tier metrics (always present)​

MetricTypeMeaning
scramdb_chunks_jit_totalcounterMorsel-chunks executed as compiled JIT machine code.
scramdb_chunks_vm_totalcounterMorsel-chunks executed by the bytecode interpreter.
scramdb_compiled_calls_totalcounterSub-region JIT invocations.

These are the counters used in the JIT scrape-delta technique described on Query Performance.

Client connections (always present)​

Every client connection holds one of [general] max_connections places while its session lives, the last superuser_reserved_connections of them kept for superusers. An idle session waits for its client without holding a thread (it is parked) and is handed back to a thread when its client sends something or a deadline passes.

MetricTypeMeaning
scramdb_max_connectionsgaugeConnection places the client listener enforces (max_connections).
scramdb_connection_places_takengaugeConnection places held by sessions.
scramdb_connection_startinggaugeConnections still negotiating SSL, sending their startup packet or authenticating.
scramdb_connection_parkedgaugeIdle sessions waiting for their client with no thread held.
scramdb_connection_worker_threadsgaugeThreads that serve client sessions, busy or cached.
scramdb_connection_idle_workersgaugeSession threads cached with no session to run.
scramdb_connection_wakes_totalcounterParked sessions handed to a thread because their client sent data or a deadline passed.
scramdb_connection_refused_too_many_totalcounterConnections refused because every connection place was taken.
scramdb_connection_refused_reserved_totalcounterConnections refused because only the places reserved for superusers were left.
scramdb_connection_refused_role_limit_totalcounterConnections refused by their role's connection limit.
scramdb_connection_refused_server_setting_totalcounterConnections refused because they asked to change a setting fixed when the server started (max_connections or superuser_reserved_connections in their startup options, FATAL 55P02).
scramdb_connection_refused_startup_setting_totalcounterConnections refused because a setting their startup packet sent (TimeZone, DateStyle or any other) had a value the setting refuses (FATAL, the setting's own code, as PostgreSQL).
scramdb_connection_hba_rejects_totalcounterConnections rejected by the host-based authentication rules when accepted.
scramdb_connection_idle_session_timeouts_totalcounterSessions closed because they stayed idle outside a transaction past idle_session_timeout (FATAL 57P05).
scramdb_connection_startup_timeouts_totalcounterConnections closed because authentication did not finish within authentication_timeout.
scramdb_connection_tls_failures_totalcounterClient TLS handshakes that failed.
scramdb_connection_messages_refused_totalcounterClient messages refused because the server could not reserve memory to receive them.
scramdb_connection_worker_start_failures_totalcounterSession threads that could not be started.
scramdb_connection_ended_client_gone_totalcounterSessions whose client closed its connection without a Terminate message.
scramdb_connection_ended_socket_error_totalcounterSessions whose client connection failed.
scramdb_connection_ended_terminated_totalcounterSessions closed by pg_terminate_backend or a server shutdown.
scramdb_connection_ended_in_transaction_totalcounterSessions that ended with a transaction open, which was rolled back (its row, table and advisory locks released with it).
scramdb_connection_startup_secondshistogramTime from accepting a connection to its session being ready for queries.

refused_too_many_total climbing means clients open more connections than max_connections allows: put a pooler in front, or raise the limit. idle_session_timeouts_total moves only when a session sets idle_session_timeout (off by default).

Statements (always present)​

A session hands each statement to a query worker and waits for its result; a statement timeout, a cancel, the transaction timeout and a lost client all end that wait at once.

MetricTypeMeaning
scramdb_statement_wait_secondshistogramTime a session waited for its statement's result from a query worker.
scramdb_statement_timeouts_totalcounterStatements cancelled because they ran past statement_timeout.
scramdb_statement_cancels_totalcounterStatements a client, a forced close or the memory watchdog cancelled while they ran.
scramdb_statement_transaction_timeouts_totalcounterSessions closed because their transaction ran past transaction_timeout during a statement.
scramdb_statement_clients_gone_totalcounterSessions closed because their client went away during a statement.
scramdb_statement_late_results_dropped_totalcounterResults of statements a session had already stopped waiting for, dropped when they arrived.
scramdb_statement_implicit_blocks_totalcounterMulti-statement queries and pipelines run as one implicit transaction, as PostgreSQL runs them.
scramdb_statement_implicit_blocks_rolled_back_totalcounterImplicit transactions rolled back because one of their statements failed.

Prepared statements (always present)​

A prepared statement is parsed once, the first time a Bind needs it; every later execution hands the engine that parse with the Bind's values in place, so a statement executed many times is never parsed again.

MetricTypeMeaning
scramdb_prepared_template_parses_totalcounterPrepared statements parsed, once each, the first time a Bind needed the parse.
scramdb_prepared_handoffs_totalcounterPrepared statement executions handed to the engine already parsed.
scramdb_prepared_handoffs_unused_totalcounterPrepared statement executions whose parse the engine did not use.
scramdb_prepared_reparses_totalcounterTimes the engine parsed an already parsed prepared statement execution again.

On a workload of prepared statements, handoffs_total grows with the executions while template_parses_total grows only with the distinct statements; handoffs_unused_total and reparses_total staying near zero confirm the parse is reused.

Plan cache (always present)​

Plans, their compiled code and plan markers are cached per statement and bounded in bytes by [optimizer] plan_cache_memory.

MetricTypeMeaning
scramdb_plan_cache_hits_totalcounterStatements that reused a cached plan.
scramdb_plan_cache_misses_totalcounterStatements that found no current cached plan and were planned.
scramdb_plan_cache_evictions_totalcounterCached plans and plan markers dropped to stay within the plan cache memory budget.
scramdb_plan_cache_bytesgaugeMemory held by cached plans, their compiled modules and plan markers.
scramdb_plan_cache_budget_bytesgaugeMemory the plan cache may hold.
scramdb_plan_cache_entriesgaugePlans and plan markers the plan cache holds.

Notifications (LISTEN and NOTIFY)​

LISTEN and NOTIFY report process-wide, always emitted:

MetricTypeMeaning
scramdb_notify_notifications_committed_totalcounterNotifications sent by committed transactions.
scramdb_notify_notifications_delivered_totalcounterNotifications written to listening clients.
scramdb_notify_batches_replicated_totalcounterCommitted transactions whose notifications were sent to the other nodes.
scramdb_notify_replication_failures_totalcounterCommitted transactions whose notifications could not be sent to the other nodes.
scramdb_notify_queue_full_totalcounterNotifications refused because the notification queue was full.
scramdb_notify_listening_sessionsgaugeSessions listening to at least one channel.
scramdb_notify_held_bytesgaugeBytes of notifications queued or waiting to be sent to a listener.

A notification's bytes are held from the execution memory budget until the last listener has received it, so held_bytes that keeps growing means a session listens and never lets its client read (a client blocked inside a long transaction receives nothing until it ends); queue_full_total counts the NOTIFY statements refused with 54000 when that budget had no room. On a cluster, replication_failures_total counts committed transactions whose notifications did not reach the other nodes: the committing client got a warning (08006) with its commit, and listeners on other nodes missed those notifications.

WAL archive metrics​

MetricTypeMeaning
scramdb_wal_archive_enabledgauge1 if WAL archival is enabled, 0 otherwise. Always emitted.
scramdb_wal_archive_errors_totalcounterArchive attempts that failed. Always emitted.
scramdb_wal_archive_segments_archived_totalcounterSegments successfully shipped to the destination. Always emitted.
scramdb_wal_archive_segments_behindgaugeSegments not yet archived. Emitted only once a real measurement exists.
scramdb_wal_archive_bytes_behindgaugeBytes not yet archived. Emitted only once a real measurement exists.
scramdb_wal_archive_last_archived_lsngaugeLSN of the last successfully archived segment. Emitted only once a real measurement exists.
scramdb_wal_archive_segments_pruned_totalcounterWAL segments removed past the archive retention window ([storage.wal.archive] retention). Always emitted.
scramdb_wal_archive_prune_failures_totalcounterArchive retention passes that failed and will retry. Always emitted; zero on a healthy archive.

The periodic checkpoint scheduler ([storage.wal.checkpoint], on by default) reports on the same endpoint, always emitted:

MetricTypeMeaning
scramdb_wal_checkpoint_interval_totalcounterCheckpoints the scheduler ran because its interval elapsed.
scramdb_wal_checkpoint_wal_bytes_totalcounterCheckpoints the scheduler ran because the WAL grew past its threshold.
scramdb_wal_checkpoint_behind_totalcounterCheckpoints run at full disk priority because the WAL grew past twice its threshold.
scramdb_wal_checkpoint_failures_totalcounterCheckpoints the scheduler started that failed.
scramdb_wal_checkpoint_threshold_bytesgaugeWAL bytes written since the last checkpoint that start the next one.
scramdb_wal_checkpoint_secondshistogramTime one scheduled checkpoint took.

The cold tier (Storage Tiering) reports on the same endpoint, always emitted; every failure counter stays at 0 on a healthy tier:

MetricTypeMeaning
scramdb_cold_segments_demoted_totalcounterSealed segments copied to cold storage.
scramdb_cold_bytes_demoted_totalcounterBytes of sealed segments copied to cold storage.
scramdb_cold_segments_evicted_totalcounterSegments whose local file was removed because cold storage holds them.
scramdb_cold_bytes_evicted_totalcounterBytes of local disk freed by segments cold storage holds.
scramdb_cold_segments_read_back_totalcounterSegments read back from cold storage to local disk.
scramdb_cold_bytes_read_back_totalcounterBytes read back from cold storage to local disk.
scramdb_cold_demotion_failures_totalcounterSegment copies to cold storage that failed and will be tried again.
scramdb_cold_read_back_failures_totalcounterSegment reads from cold storage that failed after every attempt.
scramdb_cold_read_back_retries_totalcounterSegment reads from cold storage attempted again after a failure.
scramdb_cold_read_back_mismatches_totalcounterCold storage objects that did not hold the bytes their segment's marker names.
scramdb_cold_transfer_timeouts_totalcounterTransfers to or from cold storage that ran past their deadline.

Operators past their memory (work_mem, see Memory) report their spills, always emitted:

MetricTypeMeaning
scramdb_sort_spill_runs_totalcounterSorted runs written to disk by sorts past their memory.
scramdb_sort_spill_bytes_written_totalcounterBytes of sorted runs written to disk.
scramdb_join_build_spill_partitions_totalcounterHash join build partitions written to disk to stay within their memory.

The compiled-function cache of [udf] (aot_cache_dir, aot_cache_max_bytes) reports on the same endpoint:

MetricTypeMeaning
scramdb_udf_compile_cache_bytesgaugeBytes of compiled function modules held in the on-disk compile cache.
scramdb_udf_compile_cache_limit_bytesgaugeThe most bytes the compile cache holds (aot_cache_max_bytes).
scramdb_udf_compile_cache_hits_totalcounterFunction modules loaded from the compile cache.
scramdb_udf_compile_cache_misses_totalcounterFunction modules compiled because the cache did not hold them.
scramdb_udf_compile_cache_evictions_totalcounterModules removed to keep the cache within its bound.
scramdb_udf_compile_cache_write_failures_totalcounterCompiled modules that could not be saved to the cache.
scramdb_udf_catalog_saves_totalcounterSaves of the installed modules and functions, one per INSTALL MODULE, DROP MODULE, CREATE FUNCTION and DROP FUNCTION of a module-backed language.
scramdb_udf_catalog_save_failures_totalcounterModule and function changes undone because they could not be saved (the statement failed with 58030).
scramdb_udf_catalog_modules_loaded_totalcounterInstalled modules loaded at startup.
scramdb_udf_catalog_functions_loaded_totalcounterFunctions loaded at startup.
scramdb_udf_catalog_load_failures_totalcounterSaved modules and functions that could not be loaded at startup; the server log names each.
scramdb_udf_catalog_module_files_removed_totalcounterSaved module files removed because no module uses them any more.
scramdb_udf_epoch_ticks_totalcounterTicks of the function engine's deadline clock, which ticks only while a function runs.
scramdb_udf_epoch_ticker_wakes_totalcounterTimes the function engine's deadline clock woke for a call after being idle.

The WAL ring reports on the same endpoint, always emitted. The bound is what the memory budget reserves for it: the ring, 64 times [storage.wal] flush_threshold_bytes (sized for the machine unless pinned), and its completion slots (a sixteenth of the ring). append_waits_total rising means the WAL device is not keeping up with the write rate. A stall is the log making no progress for far longer than a flush usually takes (at least a second at the default epoch): it is reported, never treated as a failure, and every commit waiting for it completes when the device answers. An interrupted log write leaves padding in its place, which every reader skips:

MetricTypeMeaning
scramdb_wal_buffer_bytesgaugeBytes of log records waiting in the WAL ring at its last flush.
scramdb_wal_buffer_bound_bytesgaugeBytes the WAL ring and its completion slots hold, as the memory budget reserves them.
scramdb_wal_append_waits_totalcounterWAL appends that waited for a flush to make room in the ring.
scramdb_wal_flush_threshold_bytesgaugeBytes of waiting log records that start a WAL flush early, as configured or sized.
scramdb_wal_flush_stalls_totalcounterTimes the WAL made no progress for far longer than a flush usually takes.
scramdb_wal_flush_stalledgaugeWAL writers whose log has made no progress for far longer than usual right now.
scramdb_wal_padding_records_totalcounterLog records whose append was interrupted and whose place was filled with padding.
scramdb_wal_flush_syncs_totalcounterData syncs of the WAL that flush rounds made, one per round that wrote log records.
scramdb_wal_rotation_syncs_totalcounterData and directory syncs of the WAL made when a file filled and the next one began.
scramdb_wal_flush_sync_secondshistogramTime one flush round's WAL data sync took.

Every sync of the WAL is counted: the sum of flush_syncs_total and rotation_syncs_total over the transactions committed in the same window is the syncs per committed transaction, and flush_sync_seconds is how long the device takes to make one flush round durable.

Catalog file writes (each table's catalog file and deleted-row record, the table list, sequences, databases and the commit timestamp mark) report on the same endpoint, always emitted. Concurrent changes to one file share one write, so coalesced_total rising while writes_total stays flat under heavy DDL or checkpointing is expected:

MetricTypeMeaning
scramdb_catalog_writes_totalcounterCatalog files written.
scramdb_catalog_write_failures_totalcounterCatalog file writes that failed. Zero on a healthy disk; each failure is returned to the statements that asked for the write.
scramdb_catalog_coalesced_totalcounterCatalog changes made durable by a write another caller did.
scramdb_catalog_write_secondshistogramTime one catalog file write took (write, sync, rename, directory sync).

The segments_behind, bytes_behind and last_archived_lsn metrics being absent from a scrape means "not measured yet," not zero. This is deliberate: the endpoint does not fabricate a 0 for a value it has not actually computed. Treat an absent metric name the same way, as unmeasured, everywhere on this endpoint, not just for WAL archive.

Object store and function IO threads​

The object store client (the s3://, gs:// and az:// destinations of the WAL archive, base backups, restores and branch staging) and the file, clock and network calls user functions make each run on a small pool of worker threads of their own. A pool starts with the first request that needs it, so a server whose destinations are local directories, or that runs no function, never starts one. Always emitted: zero until the pool starts.

MetricTypeMeaning
scramdb_object_store_io_workersgaugeWorker threads of the object store client.
scramdb_object_store_io_busy_microseconds_totalcounterTime the object store client's worker threads spent running work.
scramdb_udf_io_workersgaugeWorker threads serving functions' file, clock and network calls.
scramdb_udf_io_busy_microseconds_totalcounterTime the worker threads serving functions' IO calls spent running work.

A busy counter's rate divided by its worker gauge is that pool's utilization.

Every upload and download through the object store client is counted, and so is every HTTP read of a remote file. A request is bounded: 5 s to connect and 30 s of silence. A read that got no answer (a connection that did not open or broke, a name that did not resolve, a timeout, or a 408, 429, 500, 502, 503 or 504) is sent again, up to 10 times within 180 s, backing off from 100 ms to 15 s; an object store request keeps the same bounds. Always emitted: zero until the first request.

MetricTypeMeaning
scramdb_object_store_uploads_total / scramdb_object_store_upload_failures_totalcounterUploads to the object store, and the ones that failed.
scramdb_object_store_uploads_aborted_total / scramdb_object_store_abort_failures_totalcounterUnfinished multipart uploads aborted after a failure, and the aborts that failed (the object store keeps those parts until its own cleanup).
scramdb_object_store_uploaded_bytes_totalcounterBytes uploaded.
scramdb_object_store_downloads_total / scramdb_object_store_download_failures_totalcounterDownloads from the object store, and the ones that failed.
scramdb_object_store_downloaded_bytes_totalcounterBytes downloaded.
scramdb_object_store_upload_seconds / scramdb_object_store_download_secondshistogramTime an upload or a download took.
scramdb_remote_file_requests_total / scramdb_remote_file_request_failures_totalcounterHTTP requests for remote files, and the ones that failed after their retries.
scramdb_remote_file_request_retries_totalcounterHTTP requests for remote files sent again.
scramdb_remote_file_request_timeouts_totalcounterHTTP requests for remote files that got no answer within their bounds.

Resilience metrics (always present, all zero on a healthy process)​

MetricTypeMeaning
scramdb_parallel_join_timeout_totalcounterParallel join operations that hit their timeout.
scramdb_pool_watchdog_heals_totalcounterBuffer pool watchdog self-heal actions taken.
scramdb_pool_watchdog_heal_budget_exhausted_totalcounterTimes the watchdog's heal budget ran out.
scramdb_pool_watchdog_escalations_totalcounterTimes the watchdog escalated beyond a self-heal.

Query execution and JIT​

MetricTypeMeaning
scramdb_jit_compile_submits_totalcounterCompilation requests handed to the JIT worker. On a warm, repeating workload this should stop growing; continued growth means the compiled-artifact cache is not serving and is worth raising compiled_cache_memory for.
scramdb_chunks_jit_t1_totalcounterChunks served by tier 1 compiled code, which is the fast-to-produce tier or a cached artifact loaded from disk.
scramdb_chunks_jit_t2_totalcounterChunks served by tier 2 compiled code, the fully optimized tier.
scramdb_jit_t1_compile_secondshistogramTime to compile one module for a quick start (tier 1).
scramdb_jit_t2_compile_secondshistogramTime to compile one module fully optimized (tier 2).
scramdb_jit_jetz_load_secondshistogramTime to load one compiled module from the on-disk compile cache.
scramdb_jit_jetbc_load_secondshistogramTime to load and compile one cached intermediate module from the on-disk compile cache.
scramdb_jit_verifier_rejections_total{tier}counterCompiled kernels the code verifier rejected, labelled tier (fused, whole_function, expression_region, module). The statement still runs correctly on the interpreter; a nonzero count is a code generator defect worth reporting, not a query problem.
scramdb_udf_interpreter_warmups_totalcounterInterpreted function runtimes started, by language (python, ruby): each starts once per process, the first time a function of that language runs.
scramdb_jit_compile_memory_charges_totalcounterCompiles whose working memory was charged to the compiled-code cache budget.
scramdb_jit_compile_memory_in_flight_bytesgaugeWorking memory held by the compiles running now. It counts against compiled_cache_memory together with the cached code, and cached code is evicted to make room for it.
scramdb_jit_compiles_in_flightgaugeCompiles running now.
scramdb_jit_compile_memory_over_budget_totalcounterCompiles that started while the running compiles alone needed more than compiled_cache_memory. A compile is never refused or delayed for memory, so a growing count means the budget is too small for the compile concurrency: raise compiled_cache_memory or lower jit_compile_threads.
scramdb_jit_compile_threadsgaugeThreads of the query compile pool ([execution] jit_compile_threads, one to eight). They run on the query cores below the query threads' priority, never on the cores a cluster node reserves for its own traffic.
scramdb_jit_compile_queuedgaugeQuery compiles waiting for a compile thread.
scramdb_jit_compile_queued_bytesgaugeBytes held by the query compiles waiting for a compile thread. Bounded by compiled_cache_memory.
scramdb_jit_compile_refused_totalcounterQuery compiles not queued because the compile queue was full. The statement runs on the interpreter, never waits, and is compiled on a later run once the queue has room; a steadily growing count means compiles cannot keep up with new statement shapes: raise jit_compile_threads or compiled_cache_memory.
scramdb_vm_agg_flushes_totalcounterInterpreter-side grouped aggregate tables flushed as partial results on reaching their memory budget.
scramdb_preagg_flushes_totalcounterPer-worker aggregate tables flushed as partial results for exceeding their share of the execution budget.
scramdb_preagg_merges_on_caller_totalcounterSmall pre-aggregate merges run on the calling thread.
scramdb_preagg_merge_empty_lanes_totalcounterEmpty pre-aggregate lanes skipped at merge.
scramdb_agg_resident_bytesgaugeLive aggregate hash-table bytes: buckets, key slots, tags and string arenas.
scramdb_join_ordering_cartesian_fallbacks_totalcounterTimes the planner could not connect a join graph and fell back to emitting a cross product. Any query whose SQL does not ask for a cross join should leave this at zero; a non-zero delta is worth reporting.
scramdb_split_distinct_rewrites_totalcounterDISTINCT aggregates the planner split into two aggregation levels.
scramdb_probe_side_materializations_totalcounterOuter-join probe sides that could not stream and were written to a temporary relation first.
scramdb_join_input_materializations_totalcounterJoin inputs of any position that could not stream and were written to a temporary relation first. Falling over time on a fixed workload is the plan getting better.
scramdb_stream_topk_rows_totalcounterRows accepted by the streaming top-K sink serving ORDER BY ... LIMIT.
scramdb_topk_threshold_publishes_totalcounterTop-K thresholds published into the scan so it can skip pages that cannot beat the current best rows.
scramdb_late_fetch_rows_totalcounterRows served by late materialization: a narrow scan that fetches the remaining columns only for rows that survived.
scramdb_stream_queue_bytesgaugeBytes of worker results queued ahead of the client, bounded by backpressure.
scramdb_stream_gate_forced_admits_totalcounterResults pushed past their statement's budget because the queue was empty and there was nothing left to drain.

Startup​

MetricTypeMeaning
scramdb_startup_subsystems_pendinggaugeSubsystems a client can reach (vector indexes, the memory watchdog, the function engine) that are not ready yet. The PostgreSQL port accepts no client until this is zero.
scramdb_startup_listener_wait_microsgaugeMicroseconds the client listener waited for startup to finish before accepting clients. Zero on an ordinary boot, where every subsystem is ready before the port is opened.

No statement fails because a subsystem is still being installed: a client whose connection reaches the listener before startup has finished waits and is served once it has.

Scans and storage​

MetricTypeMeaning
scramdb_index_delete_lock_acquisitions_totalcounterBatches of entries removed from ordered and hashed indexes. Index reads never wait on these removals.
scramdb_index_delete_entries_totalcounterEntries asked to be removed from ordered and hashed indexes.
scramdb_index_read_gate_raised_totalcounterIndex key reads that held the index's writers back to finish: a read of a key that writers kept changing through a few tries pauses every change to that index until it has read the key once. Rare outside a storm of writes on one key.
scramdb_index_read_gate_writer_waits_totalcounterIndex changes that waited for such a key read to finish. Each wait lasts one read of the key.
scramdb_zone_pages_considered_totalcounterPages evaluated against zone-map predicates.
scramdb_zone_pages_pruned_totalcounterPages skipped by zone-map pruning and never read. Against considered, this is your page-pruning hit rate.
scramdb_selective_pages_read_totalcounterPages served by reading only the projected columns rather than the whole page.
scramdb_selective_bytes_saved_totalcounterBytes that column-selective reading did not have to fetch.
scramdb_selective_ranges_read_totalcounterRead requests column-selective reading issued. Divided by selective_pages_read_total it is the number of disk reads per page.
scramdb_read_pool_active_unitsgaugeUnits currently holding a read-buffer-pool slot, either waiting for concurrency or actively holding pages.
scramdb_read_pool_unit_waits_totalcounterTimes a unit waited noticeably long for a read-pool concurrency slot. This is a throughput signal, not a safety one: page acquisition itself is never queued.
scramdb_row_fetches_totalcounterFetches of rows by their address, for index probes, row re-reads and re-ranking.
scramdb_row_fetch_rows_totalcounterRows returned by fetches of rows by their address.
scramdb_row_fetch_pages_totalcounterPages read by fetches of rows by their address.
scramdb_row_fetch_page_bytes_totalcounterBytes of the pages read by fetches of rows by their address.
scramdb_row_fetch_decoded_bytes_totalcounterBytes decoded from those pages by fetches of rows by their address. Against page_bytes_total, how much of each page a fetch had to decode.
scramdb_values_out_of_line_written_totalcounterValues too long for their page written out of line, in a chain of pages of their own.
scramdb_value_chain_pages_written_totalcounterPages written to hold values stored out of line.
scramdb_values_out_of_line_bytes_written_totalcounterBytes of the values written out of line.
scramdb_value_chain_bytes_stored_totalcounterBytes stored for values written out of line, after compression. Against bytes_written_total, the compression those values got.
scramdb_values_out_of_line_read_totalcounterValues stored out of line read back.
scramdb_value_chain_pages_read_totalcounterPages read to read values stored out of line.
scramdb_values_out_of_line_bytes_read_totalcounterBytes of the values stored out of line read back.

Rows fetched by address are counted by each query worker on its own, so the counters cost nothing measurable however many workers fetch. A text, binary, vector or array value longer than fits its page is stored out of line, in a chain of pages of its row's segment, and read back only by the rows that need it; the out-of-line counters are zero on a workload with no such values.

Snapshot index reads​

An index lists only the latest version of each row. A REPEATABLE READ or SERIALIZABLE read whose snapshot is older than a change finds the version it sees through the index's record of superseded entries, instead of scanning the table.

MetricTypeMeaning
scramdb_snapshot_index_probes_versioned_totalcounterSnapshot index probes answered from the index and its record of superseded entries.
scramdb_snapshot_index_probes_scanned_totalcounterSnapshot index probes answered by scanning the table.
scramdb_retired_versions_recorded_totalcounterSuperseded index entries kept for snapshots that can still see them.
scramdb_retired_versions_released_totalcounterSuperseded index entries released by compaction or with their index.
scramdb_retired_versions_entriesgaugeSuperseded index entries currently kept for snapshots.

probes_scanned_total climbing against probes_versioned_total means snapshot reads are scanning whole tables instead of probing their index.

Read snapshot pins​

The snapshot of a statement, of a REPEATABLE READ or SERIALIZABLE transaction, of a streaming result and of a distributed fragment's local scan is pinned while it is in use, so compaction keeps every row version it can still see.

MetricTypeMeaning
scramdb_read_pins_taken_totalcounterRead snapshot pins registered.
scramdb_read_pins_released_totalcounterRead snapshot pins released.
scramdb_read_pin_redraws_totalcounterSnapshot draws that moved their pin off the floor registered before the draw.
scramdb_read_pins_heldgaugeRead snapshot pins held now.
scramdb_retention_holds_taken_totalcounterRetention holds taken for the shard groups hosted on this node.
scramdb_retention_holds_released_totalcounterRetention holds released.
scramdb_retention_holds_heldgaugeRetention holds held now for the shard groups hosted on this node.
scramdb_retention_boot_holds_armed_totalcounterStores that kept every row version at start until their shard groups were pinned.
scramdb_retention_boot_holds_released_totalcounterStart-up retention holds released.
scramdb_retention_boot_holds_heldgaugeStores keeping every row version until their shard groups are pinned.

read_pins_held that keeps growing on a steady workload means snapshots are not being released, which holds old row versions on disk; look for long-open transactions or unread result streams.

On a cluster node, each shard group hosted here also holds the row versions between the oldest read position the group can still serve and its newest commit, since a read at such a position can reach this node long after it was chosen; a quiet group holds only its last window, never what busier groups write after it. retention_holds_held follows the groups hosted here. At start a store keeps every row version until each shard group it holds rows of is hosted again or placed on another node, and it logs one line saying so; retention_boot_holds_held above zero long after start means a group was never hosted again (its restarts exhausted, or it is quarantined), and compaction on that store frees nothing until it is.

Memory admission​

Heavy statements take a memory grant before they run and queue in arrival order when the pool is full. Nothing is ever refused, so a rising queue means work is waiting, not failing.

MetricTypeMeaning
scramdb_admission_admitted_totalcounterHeavy statements that took a grant.
scramdb_admission_queued_totalcounterHeavy statements that had to wait for one, counted once each when they first park.
scramdb_admission_parked_waitersgaugeStatements parked waiting for a grant right now. Emitted only once a real measurement exists.
scramdb_admission_outstanding_grant_bytesgaugeBytes currently granted and not yet handed back.
scramdb_admission_waited_past_deadline_totalcounterStatements that waited longer than the warning threshold. Nothing is refused on that threshold; it exists so a long wait is visible.
scramdb_admission_cancelled_while_queued_totalcounterStatements whose own cancellation fired while they were still waiting.
scramdb_maintenance_memory_waits_totalcounterCREATE INDEX builds (vector indexes and their REINDEX included) and ANALYZE runs whose working memory (storage.memory.maintenance_bytes) did not fit the pool when asked, and that waited for it instead of failing. The automatic analysis a COPY or bulk INSERT asks for counts here too, and runs once the pool has room.
scramdb_analyze_on_load_waits_totalcounterAutomatic analyses of a COPY or bulk INSERT ([statistics] analyze_on_load) that waited for the load's transaction to commit, so they read the loaded rows. Growth is normal: a load inside a transaction block counts once while it waits for its COMMIT.
scramdb_admission_bypassed_totalcounterStatements served without consulting the ledger, either because they were light enough to stream or because the gate is switched off.
scramdb_admission_inherited_nested_totalcounterNested statements served inside their parent's grant instead of taking a second one. Growth is normal on a workload full of CTEs and routines.
scramdb_admission_released_at_seam_totalcounterGrants handed back cleanly at statement completion. On ordinary workloads this should climb in step with admitted_total; a gap that keeps growing means grants are being returned late, and is the first thing to check if concurrency degrades over a long run.
scramdb_admission_measured_grants_totalcounterStatements admitted with a grant sized from their measured memory.
scramdb_admission_measured_growths_totalcounterMeasured grants that had to grow while their statement ran.
scramdb_admission_heavy_concurrencygaugeHeavy statements the gate currently admits at once. Live-settable at runtime, independent of the storage.memory.heavy_concurrency config key's value at startup.

Allocator accounting​

Every byte ScramDB's own allocator hands out is counted exactly, the instant it is allocated; see Memory for how the figures below feed the budgets on that page.

MetricTypeMeaning
scramdb_heap_live_bytesgaugeThe exact process heap total: every byte the allocator has handed out and not yet freed.
scramdb_heap_reserve_bytesgaugeBytes the allocator has committed but not yet handed to a live allocation (its own retention).
scramdb_heap_overflow_cellsgaugeThreads currently sharing the heap accounting overflow cell instead of their own.
scramdb_heap_context_overflows_totalcounterCharge contexts created after the context table was full, sharing a class's permanent slot.
scramdb_heap_class_bytesgaugeLive bytes per class (statement, index_build, vector_index, background, cache, cluster, unattributed), labeled class.
scramdb_heap_escaped_bytesgaugeBytes still resident under a class whose owning statement, build or job already ended, labeled class. Steady growth names a component that keeps memory after the work that allocated it has finished.
scramdb_heap_uncharged_bytesgaugeBytes live statements hold in their allocator context past what they reserved through the execution-memory budget; the budgets already decide on the larger figure, so this shows how much work runs ahead of its reservations.
scramdb_execution_pool_allocated_bytesgaugeBytes the execution memory pool currently holds against its budget (the larger of what was reserved and what the allocator measured).
scramdb_execution_pool_budget_bytesgaugeThe execution memory pool's total budget (execution_memory_bytes).
scramdb_execution_pool_essential_breach_episodes_totalcounterTimes state a query could not give back (rows already built, mid-flight) carried the execution memory pool past its budget and its small overshoot allowance. Reported once per crossing, never once per allocation, so a healthy process reads zero or a small, stable count; a climbing count means concurrent heavy queries are routinely exceeding execution_memory_bytes together, and raising it or running fewer of them at once is the fix.

Joins over the memory budget and buffer reuse (always present)​

A hash join whose build side does not fit the execution-memory budget spills and runs in rounds, and several such joins in one probe stage share one round schedule. When the joins' keys are columns of the probe table itself, the probe side is partitioned to disk once and each pass reads only the rows its rounds can match; the query workers reuse their interpreters and join buffers across morsels. A join whose output for one batch of its probe side is larger than one output batch (a cross join, a key the build side holds many times) emits it in bounded batches, each charged to the statement's memory while it is held, and a plain query over such a join streams them to the client as they are made.

MetricTypeMeaning
scramdb_join_round_plans_totalcounterHash join builds that did not fit the execution-memory budget and ran in rounds.
scramdb_join_round_passes_totalcounterPasses over a probe side made by hash joins running in rounds.
scramdb_join_round_multi_plan_stages_totalcounterProbe stages that ran several oversized hash joins in rounds together.
scramdb_join_round_chunked_plans_totalcounterHash join builds that split one partition into several rounds to stay within the execution-memory budget.
scramdb_join_round_partitioned_probes_totalcounterProbe stages of hash joins in rounds whose probe side was partitioned to disk once instead of read on every pass.
scramdb_join_output_steps_totalcounterJoin output batches emitted by a probe whose output was larger than one batch.
scramdb_unnest_output_steps_totalcounterUnnest output batches emitted for arrays whose elements were more than one batch.
scramdb_join_output_step_bytesgaugeBytes of join and unnest output batches the query workers hold right now.
scramdb_join_output_step_peak_bytesgaugeMost bytes of join and unnest output batches the query workers held at once.
scramdb_execution_interpreters_built_totalcounterInterpreters the query workers built to run morsels on the interpreter.
scramdb_execution_join_selection_grows_totalcounterTimes a query worker grew its reused join selection buffer.
scramdb_execution_agg_rows_folded_in_runs_totalcounterGrouped aggregate rows folded into the previous row's group without a hash lookup, because their key equals the previous row's.
scramdb_execution_agg_rows_folded_in_runs_compiled_totalcounterThe same, for the rows the compiled tier folded.
scramdb_execution_agg_run_checks_switched_off_totalcounterGrouped aggregate batches whose consecutive-key check switched off for lack of runs, so input not ordered on its keys pays only a short warm-up of comparisons per batch.
scramdb_execution_call_argument_views_totalcounterBatch arguments handed to scalar functions as zero-copy views instead of copies.

The build counters stay bounded by the workers and the batch sizes, never by the rows a statement reads: a count that grows with the input points at a buffer that stopped being reused. The folded-row counters grow on input grouped or sorted by its grouping key, where most rows skip the hash table.

Spilling to disk​

A query that outgrows its memory budget writes to disk rather than failing. These count that happening.

MetricTypeMeaning
scramdb_agg_spill_partition_flushes_totalcounterAggregate partitions written to disk instead of breaching the memory budget.
scramdb_agg_spill_bytes_written_totalcounterBytes written to aggregate spill files.
scramdb_agg_spill_rows_written_totalcounterRows written to aggregate spill files.
scramdb_agg_spill_fragment_partition_flushes_totalcounterFragment partitions written to disk for the same reason.
scramdb_agg_spill_fragment_bytes_written_totalcounterBytes written to fragment spill files.
scramdb_agg_spill_fragment_rows_written_totalcounterEntries written to fragment spill files.
scramdb_spill_tape_live_bytesgaugeBytes currently resident in every live spill file. Observed, not capped: spill disk usage is deliberately unbounded so a large query completes rather than failing. Alert on it if the volume is small.
scramdb_fragment_spill_mgr_init_failures_totalcounterTimes the fragment spill directory could not be initialized and was retried. Zero on a healthy process; non-zero points at temp_dir permissions or space.
scramdb_final_agg_spill_mgr_init_failures_totalcounterThe same, for the final aggregate spill directory.

Bulk load​

MetricTypeMeaning
scramdb_copy_one_pass_chunks_totalcounterChunks processed by the single-read COPY parser. Zero means that path never engaged, either because it is switched off or because no eligible COPY has run.
scramdb_copy_one_pass_reparses_totalcounterChunks that had to be parsed a second time because the parser's guess about where quoting stood at the chunk boundary was wrong. Against chunks_total this is the miss rate; a file with heavy embedded quoting raises it.
scramdb_copy_one_pass_reparse_bytes_totalcounterBytes re-read by those second passes, tracked separately so load throughput is never overstated.
scramdb_copy_staged_bytesgaugeBytes of staged COPY rows and key runs held on this node.
scramdb_copy_stages_opengaugeCOPY stages holding rows that are not committed yet: a node-local COPY's, and on a cluster node also every replica's stage of a cluster COPY and the coordinating node's stage of the COPY's key checks.
scramdb_copy_staged_rows_totalcounterRows COPY staged.
scramdb_copy_stages_attached_totalcounterCOPY stages committed into their tables.
scramdb_copy_stages_dropped_totalcounterCOPY stages discarded without a commit.
scramdb_copy_staging_refusals_totalcounterCOPY batches refused because the staging area was full. Each one failed its COPY with the disk-full error, leaving nothing behind.
scramdb_copy_stages_recovered_totalcounterCOPY stages a restart found and freed.
scramdb_copy_stages_restored_totalcounterCommitted COPY stages a restore or a branch rebuilt from the WAL.
scramdb_copy_attach_secondshistogramTime to attach one COPY stage to its table at commit.

A COPY writes its rows into staged segments outside the table, invisible to every reader, and attaches them to the table in one step when its transaction commits; a rollback or a crash before the commit frees them whole. copy_staged_bytes is the disk those uncommitted rows hold now. A COPY into a cluster table stages its rows on every replica of every shard group it writes and commits them together (see COPY into a distributed table); copy_stages_open above zero while no COPY is running means a stage was left behind, which a group leader collects after [cluster.dilith] copy_stage_idle_timeout (scramdb_cluster_dilith_copy_stages_collected_total).

Import (COPY ... FROM STDIN)​

MetricTypeMeaning
scramdb_copy_in_started_totalcounterCOPY FROM STDIN statements that started receiving rows from a client.
scramdb_copy_in_completed_totalcounterCOPY FROM STDIN statements that loaded their rows.
scramdb_copy_in_failed_totalcounterCOPY FROM STDIN statements that failed, including those the client aborted.
scramdb_copy_in_client_aborts_totalcounterCOPY FROM STDIN statements the client ended with CopyFail.
scramdb_copy_in_bytes_totalcounterBytes of COPY FROM STDIN input handed to loads.
scramdb_copy_in_slices_totalcounterSlices of COPY FROM STDIN input handed to loads.
scramdb_copy_in_client_waits_totalcounterTimes a COPY FROM STDIN's input waited for its load to take the slices before it.

An import queues at most [execution] copy_pipeline_depth slices of input for its load and makes the client wait while the queue is full, so copy_in_client_waits_total climbing against copy_in_started_total means the load is slower than the client sends; the server is not buffering the difference.

Export (COPY ... TO STDOUT)​

MetricTypeMeaning
scramdb_copy_out_streams_totalcounterCOPY TO STDOUT statements that started streaming rows to a client.
scramdb_copy_out_rows_totalcounterRows COPY TO STDOUT statements streamed to clients.
scramdb_copy_out_bytes_totalcounterBytes of COPY TO STDOUT output streamed to clients.
scramdb_copy_out_failed_totalcounterCOPY TO STDOUT statements that ended with an error after they started streaming.
scramdb_copy_out_client_waits_totalcounterTimes a COPY TO STDOUT waited for its client to take rows before rendering more.
scramdb_copy_out_secondshistogramTime from a COPY TO STDOUT's CopyOutResponse to its CopyDone.

An export holds [execution] copy_pipeline_depth buffers and waits for the client when they are full, so copy_out_client_waits_total climbing against copy_out_streams_total means clients read slower than the server renders; the export is not buffering the difference.

Foreign keys​

MetricTypeMeaning
scramdb_fk_referencing_keys_checked_totalcounterReferencing keys checked against the table they reference.
scramdb_fk_referenced_keys_checked_totalcounterReferenced keys that left or changed and were checked for rows still naming them.
scramdb_fk_violations_totalcounterStatements and commits refused because a foreign key was violated.
scramdb_fk_actions_totalcounterReferential actions (CASCADE, SET NULL, SET DEFAULT) run on referencing tables.
scramdb_fk_deferred_checks_totalcounterForeign key checks deferred to the transaction's commit point.
scramdb_fk_key_shares_totalcounterReferenced keys a referencing check held shared while it ran, so a concurrent delete of the key waits for it.
scramdb_fk_key_share_waits_totalcounterReferencing checks that waited for a transaction removing the key they reference.
scramdb_fk_leaving_key_locks_totalcounterReferenced keys held exclusively because they left or changed.
scramdb_fk_removed_after_snapshot_totalcounterReferencing checks refused with 40001 because a key their snapshot held was removed by a later commit.
scramdb_fk_probe_secondshistogramTime one foreign key probe statement took.

Deep diagnostics​

These are emitted but are not part of the supported surface, and their names and meanings can change between releases. They are here so a number you see on the endpoint is never unexplained; do not build alerts on them.

MetricTypeMeaning
scramdb_header_sidecar_loads_totalcounterSegments whose page headers were loaded entirely from their persisted sidecar file: one sequential read instead of one per page.
scramdb_header_sidecar_partial_loads_totalcounterSidecars that covered only part of what was asked for, the ordinary case for a table still growing.
scramdb_header_sidecar_fallbacks_totalcounterSidecar reads that fell back to walking page headers because the file was absent, stale or corrupt.
scramdb_header_sidecar_writes_totalcounterHeader sidecars written.
scramdb_header_sidecar_write_errors_totalcounterHeader sidecar writes that failed. The segment is walked next time instead.
scramdb_header_prefix_pages_walked_totalcounterPages read by the per-page header walk. A warm store should stop moving it.
scramdb_page_zero_bytes_full_totalcounterBytes zeroed over a whole page buffer on release.
scramdb_page_zero_bytes_partial_totalcounterBytes zeroed over only the recorded written region. On a warm selective scan this should dominate the full counter.
scramdb_scan_prune_nanos_totalcounterCore-nanoseconds spent choosing which pages to read. Summed across workers, so it exceeds wall-clock time on a parallel scan.
scramdb_scan_acquire_nanos_totalcounterCore-nanoseconds spent acquiring pages, pinning them when resident and reading them when not.
scramdb_scan_decode_nanos_totalcounterCore-nanoseconds spent turning page bytes into chunks.
scramdb_agg_fold_rows_totalcounterPartial aggregate rows consumed by the final fold.
scramdb_agg_fold_chunks_parallel_totalcounterFold chunks large enough to spread across cores.
scramdb_agg_fold_chunks_in_place_totalcounterFold chunks small enough to fold on one core.
scramdb_agg_fold_distinct_values_offered_totalcounterDISTINCT values offered to a fold's deduplication set.

Licensing metrics​

MetricTypeMeaning
scramdb_license_validgauge1 if the active license is valid.
scramdb_license_editiongauge/labelActive license edition.
scramdb_license_node_countgaugeNodes currently counted against the license.
scramdb_license_max_nodesgaugeNode limit the license allows.
scramdb_license_read_only_activegauge1 if the license has forced the deployment read-only.
scramdb_license_days_of_term_remaininggaugeEmitted only for a term-limited license.
scramdb_license_data_bytesgaugeEmitted only for a license with a data-size cap.
scramdb_license_data_capgaugeEmitted only for a license with a data-size cap.

Cluster mode​

Running as a cluster node, additional scramdb_cluster_* metrics are appended to the same /metrics response after the single-node metrics above. Every failure mode below has a counter: a stalled compaction, a quarantined group, a dropped partition frame. Alerting on a ScramDB cluster means reading a number, never grepping logs for a string match.

Unlike the WAL archive metrics above, every scramdb_cluster_* metric is always present once [cluster] is configured, whether or not anything interesting has happened yet. A fresh cluster reports every counter at 0 and every gauge at its idle value; there's no "not measured yet" gap to account for in this family.

Peer transport​

MetricTypeMeaning
scramdb_cluster_peer_connects_totalcounterPeer connections established over the cluster transport.
scramdb_cluster_peer_disconnects_totalcounterPeer connections closed.
scramdb_cluster_reconnect_attempts_totalcounterReconnect attempts after a lost peer connection.
scramdb_cluster_frames_sent_totalcounterFrames sent to peers.
scramdb_cluster_frames_received_totalcounterFrames received from peers.
scramdb_cluster_bytes_sent_totalcounterBytes sent to peers.
scramdb_cluster_bytes_received_totalcounterBytes received from peers.
scramdb_cluster_send_queue_bytesgaugeOutbound send-queue occupancy across all peer connections, right now. Bounded by send_queue_bytes in [cluster] config (default 64MB) per connection; a value pinned near that bound is backpressure, not idle capacity.
scramdb_cluster_partition_drops_totalcounterFrames silently dropped by an injected network partition, outbound and inbound combined. Zero on every normal deployment; nonzero means a chaos/test partition is, or was, active.
scramdb_cluster_fragment_boot_refusals_totalcounterFragment requests from other nodes answered with a refusal because this node was still starting.
scramdb_cluster_forward_boot_refusals_totalcounterRequests other nodes forwarded to this node for the metadata group (a DDL, a lock, a node number claim, a member removal) answered as not made because this node was still starting: the sender asks the next node at once instead of waiting out its forward timeout.
scramdb_cluster_handshake_latency_secondshistogramPeer handshake latency. 11 buckets from 1ms to 5s (le="0.001" through le="5.000") plus +Inf.
scramdb_cluster_connections_closed_totalcounterPeer connections closed, by cause: ping_deadline (no byte from the peer within its dead-peer deadline, the stalled-process case), tcp_timeout (the kernel's user timeout or keepalive gave up), peer_closed, local (shutdown or a duplicate connection), io_error, protocol_error.
scramdb_cluster_frames_dropped_on_close_totalcounterFrames still queued on a peer connection when it closed, dropped and reported to the subsystems that sent them.
scramdb_cluster_peer_events_totalcounterPeer connection events published to the cluster's subsystems, by kind: connected and lost.
scramdb_cluster_peer_event_lag_totalcounterTimes a subsystem fell behind the connection event channel and treated every peer as lost once. Zero in normal operation.
scramdb_cluster_liveness_pings_sent_totalcounterLiveness pings sent on peer connections that had sent nothing for ping_interval.
scramdb_cluster_liveness_ping_rtt_secondshistogramRound trip of a liveness ping, from its send to its answer. 18 buckets from 10µs to 10s plus +Inf.
scramdb_cluster_dead_peer_deadline_connectionsgaugeOpen peer connections by how their dead-peer deadline is set, source: floor (the adaptive deadline at dead_peer_timeout_min), measured (lifted by a long round trip), ceiling (held at dead_peer_timeout_max), pinned (a fixed dead_peer_timeout).
scramdb_cluster_dead_peer_deadline_secondshistogramEvery dead-peer deadline put in force, at a connection's start and whenever its measured round trip moves it.
scramdb_cluster_socket_option_failures_totalcounterSocket options (TCP_NODELAY, keepalive, user timeout) the kernel refused on a peer connection. Zero on Linux; nonzero means dead-peer detection relies on the ping deadline alone.
scramdb_cluster_peers_rememberedgaugePeers whose connection closed within the grace period and whose handshake details are kept.
scramdb_cluster_peers_remembered_expired_totalcounterRemembered peers dropped after staying disconnected past the grace period.
scramdb_cluster_peers_left_totalcounterPeers forgotten because they left the cluster membership.
scramdb_cluster_dial_backsgaugePeers this node is dialing back because this node must open their connection.
scramdb_cluster_dial_backs_started_totalcounterTimes this node started dialing a peer back.
scramdb_cluster_dial_backs_ended_totalcounterTimes this node stopped dialing a peer back.
scramdb_cluster_peer_identity_refusals_totalcounterPeer connections refused because the certificate does not carry the node name the peer announced. Nonzero means a node's certificate is missing its name; see TLS between cluster nodes.
scramdb_cluster_duplicate_connections_refused_totalcounterPeer connections refused because this node already accepted one from the same peer on the same lane. A peer that opens a lane twice at once (two seed addresses for one node, or a reconnect racing the connection it replaces) keeps the first and drops the second, so a few are normal; a steady climb means a peer keeps reopening a lane it already holds.
scramdb_cluster_transport_threadsgaugeThreads the cluster transport runs for its listeners, dials, handshakes and peer connections.
scramdb_cluster_transport_thread_start_failures_totalcounterThreads the cluster transport could not start.

During an injected partition, partition_drops_total also counts the dropped liveness pings and every message that was already queued to a cut peer when the partition started (a message already partly sent is finished once the partition heals), and the connection to each cut peer is closed by the ping deadline and re-established every couple of seconds until the partition heals, so peer_connects_total, peer_disconnects_total and reconnect_attempts_total climb with it.

Traffic lanes​

Every pair of nodes talks over three connections, one per traffic lane, each on its own port: control (membership, consensus, the commit protocol, forwarded DDL and lock calls, shuffle fetches), interactive (replies to forwarded point reads) and bulk (query data exchange, fragment dispatch with its broadcast join side, snapshot transfer, the rows a cluster COPY stages). A message larger than [cluster.transport] chunk_size travels as chunks, and a lane's connection takes its streams in turn a chunk at a time. Every stream has its own window of its lane, so a subsystem that stops reading holds only its window and stops only its own senders.

MetricTypeMeaning
scramdb_cluster_lane_bytes_sent_totalcounterBytes written to peers, by lane.
scramdb_cluster_lane_frames_sent_totalcounterFrames written to peers, by lane.
scramdb_cluster_lane_bytes_received_totalcounterBytes read from peers, by lane.
scramdb_cluster_lane_frames_received_totalcounterFrames read from peers, by lane.
scramdb_cluster_lane_chunks_sent_totalcounterChunks of large messages written to peers, by lane.
scramdb_cluster_lane_chunks_received_totalcounterChunks of large messages read from peers, by lane.
scramdb_cluster_lane_messages_reassembled_totalcounterLarge messages reassembled from chunks, by lane.
scramdb_cluster_lane_credit_frames_sent_totalcounterFrames that returned stream credit to a peer, by lane.
scramdb_cluster_lane_connectionsgaugeOpen peer connections, by lane.
scramdb_cluster_lane_older_protocol_connectionsgaugeOpen peer connections that speak an older cluster protocol version, by traffic lane.
scramdb_cluster_lane_protocol_fallbacks_totalcounterDials made again with the older cluster protocol version a peer asked for, by traffic lane.
scramdb_cluster_lane_stream_window_bytesgaugeReceive window granted to each stream on the newest connection, by lane: the lane's size shared by the peers then connected.
scramdb_cluster_lane_queued_bytesgaugeBytes held by each traffic queue, by lane and direction: send (queued to send) or receive (received and not yet consumed).
scramdb_cluster_lane_capacity_bytesgaugeSize of each traffic queue, by lane and direction.
scramdb_cluster_lane_backpressure_totalcounterSends refused because a traffic queue or a stream window was full, by lane and stream (membership, metadata_consensus, shard_consensus, commit_protocol, ddl_forward, lock_forward, shuffle_fetch, point_read, exchange, fragment_dispatch, snapshot_transfer, copy_staging, other).
scramdb_cluster_lane_drain_waits_totalcounterSends that waited for a full traffic queue to drain instead of failing.
scramdb_cluster_lane_drain_wait_secondshistogramTime a send waited for a full traffic queue to drain.
scramdb_cluster_lane_connection_threadsgaugeReader and writer threads of peer connections, by lane: two per open connection.

backpressure_total on the bulk lane is flow control doing its job under a large query. On the control lane it should stay at zero: a nonzero shard_consensus or metadata_consensus count means consensus traffic found its queue full, and each such event is also logged once.

Raft group health and recovery​

MetricTypeMeaning
scramdb_cluster_group_quarantines_totalcounterRaft group storage directories quarantined: at boot for deeper corruption (a bad header, a mid-prefix CRC failure, an undecodable hard state), or at runtime when a group stopped for a reason a restart cannot fix. Does not count a benign torn-tail truncation; that's a row below.
scramdb_cluster_quarantined_groupsgaugeGroups currently quarantined, awaiting a membership-replace rejoin. This is the line to alert on.
scramdb_cluster_group_rejoins_totalcounterQuarantined groups that completed their membership-replace rejoin.
scramdb_cluster_term_changes_totalcounterMetadata group (group0) raft term changes this node has observed: an election, a term bump from a higher-term message, or a step-down. Node-local, not a cluster-wide total: each node counts its own view of the same election.
scramdb_cluster_group_actor_restarts_totalcounterShard groups this node restarted in place after a transient stop, within a bounded restart budget; one that keeps stopping is quarantined instead.
scramdb_cluster_torn_tail_truncations_totalcounterBenign torn-tail truncations at raft log open: the normal crash-recovery path, made visible instead of silent.

scramdb_cluster_quarantined_groups is restart-honest. It is never incremented or decremented as events happen; every host-reconcile round recomputes it from scratch, from the actual set of groups gated right now, and overwrites the gauge with that count. A crash that loses an in-flight update, or a process restart mid-recovery, can't leave this gauge stuck on a stale nonzero reading or silently reset to a wrong zero: the very next reconcile round always reflects what's really quarantined at that moment, never a running total of past events.

Consensus threads​

Every replication group and the metadata group of a node run on one pool of consensus threads ([cluster.consensus] threads), never a thread per group. A group runs one step at a time and hands its thread back, so a slow group holds one thread and never another group.

MetricTypeMeaning
scramdb_cluster_consensus_pool_threadsgaugeThreads of the pool that drives the consensus groups.
scramdb_cluster_consensus_pool_queuedgaugeConsensus groups and cluster tasks waiting for a consensus thread. Near zero on a healthy node; a value that stays high means the threads cannot keep up.
scramdb_cluster_consensus_pool_panics_totalcounterConsensus groups and cluster tasks stopped by a panic. Zero on a healthy node; the log names the cause.
scramdb_cluster_consensus_mailbox_refused_totalcounterMessages a consensus group refused because its queue was full or closed.

Exchange threads​

The distributed query work other nodes send a node (their fragments, the rows arriving, and the results sent back) runs on its own pool of exchange threads ([cluster.exchange] threads), never on the consensus threads and never on the threads that run queries.

MetricTypeMeaning
scramdb_cluster_exchange_pool_threadsgaugeThreads of the pool that runs distributed query work sent by other nodes.
scramdb_cluster_exchange_pool_queuedgaugeDistributed query tasks waiting for an exchange thread. Near zero on a healthy node; a value that stays high means the threads cannot keep up.
scramdb_cluster_exchange_pool_panics_totalcounterDistributed query tasks stopped by a panic. Zero on a healthy node; the log names the cause.

Cluster services​

Requests other nodes send this node's cluster services (lock and DDL forwarding, snapshot transfer, query fragments and shuffle fetches) are served on engine threads, each service with its own bound on the requests it serves at once.

MetricTypeMeaning
scramdb_cluster_stream_requests_totalcounterRequests from other nodes this node's cluster services started serving.
scramdb_cluster_stream_requests_in_flightgaugeRequests from other nodes this node's cluster services are serving now.
scramdb_cluster_stream_requests_waited_totalcounterRequests that waited for a free place because their service was at its bound.
scramdb_cluster_stream_frames_refused_totalcounterFrames dropped because their service had stopped or its queue was full.

Cluster-wide LOCK TABLE grants and transaction advisory locks are released at the end of the transaction or the session that holds them, a session that ends with its transaction open included:

MetricTypeMeaning
scramdb_cluster_lock_grants_released_totalcounterCluster-wide lock grants released at the end of a transaction or session.
scramdb_cluster_lock_grant_release_failures_totalcounterCluster-wide lock grants whose release at the end of a transaction or session failed.
scramdb_cluster_lock_detached_releases_totalcounterSessions that ended holding cluster-wide lock grants, released after they closed.

Group commit and durability​

MetricTypeMeaning
scramdb_cluster_group_commit_ops_totalcounterRaft ops served, summed across every group-commit wake-up. Divide by group_commit_wakeups_total for the achieved batch size. A lone proposal (nothing to batch with) contributes 1 op and 0 wake-ups, so an idle system reports no batching instead of a fake 1.0.
scramdb_cluster_group_commit_wakeups_totalcounterWake-ups that served more than one op, i.e. wake-ups that actually batched.
scramdb_cluster_group_commit_max_batchgaugeThe largest batch any wake-up has served so far, a high-water mark. Pinned at your configured cap says "raise the cap"; comfortably under it says the batching window is already draining faster than proposals arrive.
scramdb_cluster_durability_flushes_totalcounterDurability flushes, one per fsync. Divide into group_commit_ops_total for fsyncs-per-commit, the number batching exists to lower.
scramdb_cluster_log_compactions_totalcounterRaft log compactions performed.
scramdb_cluster_log_entries_compacted_totalcounterRaft log entries reclaimed by compaction, summed.

Zero log_compactions_total is normal on a cluster whose log never reaches the compaction trigger. A flat zero on a cluster that's otherwise busy (group_commit_ops_total climbing) is not normal: compaction is waiting, either for a state snapshot that keeps failing (state_snapshot_failures_total or state_snapshots_abandoned_total climbing: a full disk, or a [cluster.log_compaction] memory_bytes budget too small for the group's state) or for a replica that is behind but still inside the lag limit. A replica lagging further than replica_lag_limit_entries no longer holds the log: when it returns it is repaired from a state snapshot, so a dead replica bounds the log instead of growing it forever. The waiting case also logs on its own, but the metric is what you'd wire an alert to; see Alerting on cluster metrics below.

Shared raft log​

Every shard group a node hosts writes its replication log into the node's one shared raft log: one writer makes a whole batch of appends from every group durable with one data sync, so a transaction that touches several groups costs the node one sync, not one per group. Its segment files are reused a whole segment at a time once nothing in them is needed; [cluster.consensus] log_segment_size sets their size (by default a share of the data volume). The metadata group keeps its own log.

MetricTypeMeaning
scramdb_cluster_shared_log_drains_totalcounterBatches the writer of the node's shared raft log wrote, each ending in one sync.
scramdb_cluster_shared_log_syncs_totalcounterData syncs of the node's shared raft log.
scramdb_cluster_shared_log_group_syncs_totalcounterGroups whose writes a sync of the node's shared raft log covered, summed over its syncs. Divided by syncs_total, the mean groups one sync made durable.
scramdb_cluster_shared_log_records_totalcounterRecords written to the node's shared raft log.
scramdb_cluster_shared_log_bytes_written_totalcounterBytes written to the node's shared raft log.
scramdb_cluster_shared_log_copied_bytes_totalcounterEntry payload bytes copied into write buffers of the node's shared raft log; larger payloads are written in place.
scramdb_cluster_shared_log_truncations_totalcounterTruncate records written for a replica's conflicting log suffix.
scramdb_cluster_shared_log_segments_created_totalcounterSegment files the node's shared raft log started.
scramdb_cluster_shared_log_segments_removed_totalcounterSegment files of the node's shared raft log removed once nothing in them was needed.
scramdb_cluster_shared_log_compactions_requested_totalcounterGroups asked to compact their log because they held the oldest segment of the node's shared raft log.
scramdb_cluster_shared_log_relocated_records_totalcounterRecords rewritten at the head of the node's shared raft log so an old segment could be removed.
scramdb_cluster_shared_log_relocated_bytes_totalcounterBytes rewritten at the head of the node's shared raft log so an old segment could be removed.
scramdb_cluster_shared_log_sync_failures_totalcounterFailed writes or syncs of the node's shared raft log; the first one stops every group on the node.
scramdb_cluster_shared_log_torn_bytes_totalcounterBytes of an unfinished write cut from the end of the node's shared raft log at start.
scramdb_cluster_shared_log_migrated_groups_totalcounterGroups whose own raft log was moved into the node's shared raft log.
scramdb_cluster_shared_log_segmentsgaugeSegment files of the node's shared raft log.
scramdb_cluster_shared_log_bytesgaugeBytes of the node's shared raft log on disk.
scramdb_cluster_shared_log_live_bytesgaugeBytes of the node's shared raft log still needed by a group.
scramdb_cluster_shared_log_groupsgaugeGroups with records in the node's shared raft log or a store open on it.
scramdb_cluster_shared_log_stoppedgauge1 once a failed write or sync stopped the node's shared raft log, else 0.
scramdb_cluster_shared_log_segment_bytesgaugeThe segment size of the node's shared raft log in force.
scramdb_cluster_shared_log_scan_millisecondsgaugeHow long the node's shared raft log took to read its segments at start.
scramdb_cluster_shared_log_sync_secondshistogramTime one data sync of the node's shared raft log took.
scramdb_cluster_shared_log_drain_secondshistogramTime from the start of a batch's write to its sync returning.

shared_log_stopped at 1 is the one to page on: a write or sync of the log failed, every shard group on the node stopped rather than trust a disk whose state is no longer known, and the node must be restarted (nothing is retried). shared_log_bytes far above shared_log_live_bytes while compactions_requested_total climbs means an idle group holds the oldest segment; the log asks it to compact and then moves its few live records forward, which relocated_bytes_total counts.

State snapshots and log compaction​

A data group compacts its replication log behind a snapshot of its commit state: the open transactions' locks and staged writes, the recent versions of every row it wrote, and its read and commit watermarks, taken at one applied log position. A restart restores the newest snapshot and replays only the log above it; a replica that fell behind the compaction point installs the leader's snapshot and rebuilds the rows it missed from it instead of copying the whole group. The cluster metadata log compacts the same way. When a group compacts, how far a replica may lag and how much memory snapshots may hold are set under [cluster.log_compaction]; each count adapts to the group's own snapshot and entry sizes unless you pin it.

MetricTypeMeaning
scramdb_cluster_state_snapshots_taken_totalcounterState snapshots taken for a log compaction.
scramdb_cluster_state_snapshot_failures_totalcounterState snapshots that could not be taken (an I/O error, a full disk).
scramdb_cluster_state_snapshots_abandoned_totalcounterState snapshots given up to stay within their memory budget.
scramdb_cluster_state_snapshot_last_bytesgaugeSize of the most recent state snapshot taken on this node.
scramdb_cluster_state_snapshot_bytes_totalcounterBytes of every state snapshot taken.
scramdb_cluster_state_snapshot_duration_secondshistogramTime from the start of a state snapshot to its completion. A snapshot is written in small steps between applied batches, so the group keeps applying while it runs.
scramdb_cluster_state_snapshot_step_secondshistogramTime one of those steps held a group's apply.
scramdb_cluster_state_snapshot_memory_bytesgaugeMemory held right now by state snapshots being taken, sent or received.
scramdb_cluster_state_snapshot_memory_limit_bytesgaugeMemory they may hold at once (memory_bytes, or the derived default).
scramdb_cluster_state_snapshot_restores_totalcounterGroups restored from a state snapshot at startup.
scramdb_cluster_state_snapshot_missing_totalcounterGroups found at startup with a compacted log and no state snapshot to restore from. Zero in a healthy deployment; anything else means that replica needs a full copy.
scramdb_cluster_state_snapshot_restore_secondshistogramTime to restore one group's state at startup.
scramdb_cluster_log_entries_replayed_totalcounterLog entries replayed at startup on top of the restored state.
scramdb_cluster_log_torn_tail_replays_totalcounterGroups found at startup with their log cut below what they had applied (a crash, or a disk that lost part of the log). The group takes the cut entries from its peers again and replays them without applying their rows a second time.
scramdb_cluster_log_compactions_awaiting_snapshot_totalcounterLog compactions held back until a state snapshot covers the open transactions.
scramdb_cluster_log_compactions_past_lagging_replica_totalcounterLog compactions that passed a replica lagging beyond the lag limit. That replica is repaired from the snapshot when it returns.
scramdb_cluster_state_snapshot_installs_totalcounterState snapshots from a leader installed on this node.
scramdb_cluster_state_snapshot_installs_incomplete_totalcounterSnapshot installs whose skipped rows come from a full copy (every install: a snapshot carries no row).
scramdb_cluster_state_snapshot_fetches_totalcounterState snapshots fetched from a peer, in bounded chunks.
scramdb_cluster_state_snapshot_fetch_failures_totalcounterState snapshot fetches from a peer that failed; the fetch retries with a growing pause.
scramdb_cluster_state_snapshot_fetched_bytes_totalcounterBytes of state snapshots fetched from peers.
scramdb_cluster_state_snapshot_chunks_served_totalcounterState snapshot chunks served to peers.
scramdb_cluster_replica_copies_started_totalcounterFull copies of a shard group's rows this node started after a snapshot install.
scramdb_cluster_replica_copies_completed_totalcounterFull copies of a shard group's rows this node completed.
scramdb_cluster_replicas_awaiting_rowsgaugeShard group replicas on this node waiting for a full copy; they serve no read. A value that stays above zero means a copy cannot finish: check replica_copy_attempt_failures_total here and replica_copy_build_failures_total on the group's voters.
scramdb_cluster_replica_copy_attempt_failures_totalcounterAttempts to copy a shard group's rows from a peer that failed and were retried.
scramdb_cluster_replica_copy_rows_totalcounterRows this node received in full copies of shard groups.
scramdb_cluster_replica_copy_bytes_totalcounterBytes this node received in full copies of shard groups.
scramdb_cluster_replica_copy_secondshistogramTime from a snapshot install to its full copy being complete.
scramdb_cluster_replica_copies_built_totalcounterFull copies of a shard group's rows this node built for a peer.
scramdb_cluster_replica_copy_build_failures_totalcounterFull copies of a shard group's rows this node could not build for a peer.
scramdb_cluster_replica_copy_chunks_served_totalcounterChunks of full copies this node served to peers.
scramdb_cluster_metadata_log_compactions_totalcounterCluster metadata log compactions.
scramdb_cluster_metadata_state_restores_totalcounterCluster metadata states restored or installed from a state snapshot.
scramdb_cluster_log_compaction_trigger_entriesgaugeRetained log entries at which a data group last decided to compact.
scramdb_cluster_replica_lag_limit_entriesgaugeLog entries a replica may lag before it stops holding the leader's log.
scramdb_cluster_log_compaction_trigger_sourcegaugeWhy the trigger has its value: 0 configured, 1 the floor (nothing measured yet, or measured below it), 2 measured from the group's snapshot and entry sizes, 3 the ceiling.

Cluster archive​

With [storage.wal.archive] enabled, each shard group's leader ships the group's committed log, and the metadata group's leader the metadata log, to the archive destination, for cluster restore. Each segment lands under a key only one leader can create, so a leader that lost its leadership can never overwrite what the new one shipped, and once it finds the group archived in a newer term it stops archiving that group.

MetricTypeMeaning
scramdb_cluster_archive_segments_totalcounterLog segments this node shipped to the cluster archive.
scramdb_cluster_archive_entries_totalcounterLog entries this node shipped to the cluster archive.
scramdb_cluster_archive_bytes_totalcounterBytes of log segments this node shipped to the cluster archive.
scramdb_cluster_archive_fenced_totalcounterLog segments refused because another leader had already archived that log position.
scramdb_cluster_archive_superseded_totalcounterGroups this node stopped archiving because a leader of a newer term archives them.
scramdb_cluster_archive_bases_totalcounterGroup bases (a state image and a full copy of the rows) this node shipped.
scramdb_cluster_archive_images_totalcounterGroup state images this node shipped alone, between bases.
scramdb_cluster_archive_gaps_totalcounterTimes a group's log was compacted past the archive and no base could be shipped.
scramdb_cluster_archive_errors_totalcounterFailed attempts to write to the cluster archive.
scramdb_cluster_archive_pruned_totalcounterArchive objects the retention pass removed.
scramdb_cluster_archive_time_samples_totalcounterTime index samples this node shipped to the cluster archive.
scramdb_cluster_archive_lag_entriesgaugeCommitted log entries of the groups this node leads that are not archived yet.
scramdb_cluster_archive_groupsgaugeGroups whose log this node ships now.
scramdb_cluster_archive_ship_secondshistogramTime to ship one log segment to the cluster archive.

A few fenced_total around a leader change are expected: the old leader's last attempt is refused because the new leader already archived that position. gaps_total above zero means a group's history has a hole a restore cannot cross (the group compacted its log past what was archived and no base could be shipped): a restore to a point inside that hole is refused, and points after the group's next base restore again. lag_entries growing without bound means the destination cannot keep up.

Distributed transaction commit​

What the commit protocol of distributed transactions ([cluster.dilith] in the config file) did on this node.

MetricTypeMeaning
scramdb_cluster_dilith_commits_totalcounterTransactions committed.
scramdb_cluster_dilith_aborts_totalcounterTransaction attempts aborted, by reason: conflict (another transaction held a pending write or a read fence the vote needed), stale_read (a version above a read position, or a read below the retention floor), outdated_vote (older than the group's memory of decisions), wait_die (the younger of two transactions on one row), recovery (a recovery found a group that never voted).
scramdb_cluster_dilith_commit_latency_secondshistogramTime from a commit request to its answer.
scramdb_cluster_dilith_vote_round_secondshistogramTime from sending a transaction's votes to its last vote answer.
scramdb_cluster_dilith_votes_appended_totalcounterVote entries appended by group leaders.
scramdb_cluster_dilith_precheck_refusals_totalcounterVotes a group leader refused before appending anything.
scramdb_cluster_dilith_parks_totalcounterRequests parked at a group behind an undecided transaction.
scramdb_cluster_dilith_write_free_certified_totalcounterTransactions that wrote nothing certified without a log entry.
scramdb_cluster_dilith_fences_appended_totalcounterFence entries appended for analytical cuts.
scramdb_cluster_dilith_fences_avoided_totalcounterAnalytical cuts a group already covered without a fence entry.
scramdb_cluster_dilith_recoveries_started_totalcounterRecoveries started for undecided transactions.
scramdb_cluster_dilith_recoveries_decided_totalcounterRecoveries that reached a verdict.
scramdb_cluster_dilith_recoveries_behind_waiter_totalcounterRecoveries started by a request that waited too long behind an undecided transaction.
scramdb_cluster_dilith_recoveries_abandoned_vote_totalcounterRecoveries started for votes pending too long with nobody waiting.
scramdb_cluster_dilith_recoveries_handed_over_totalcounterRecoveries handed to the group leader that holds the pending vote.
scramdb_cluster_dilith_recoveries_from_record_totalcounterRecoveries decided from the leader's own record of the transaction.
scramdb_cluster_dilith_status_probes_totalcounterStatus queries for transactions whose outcome a client did not receive.
scramdb_cluster_dilith_sweeps_totalcounterSweep entries appended to decide transactions that never voted at a group.
scramdb_cluster_dilith_sweeps_refused_totalcounterSweeps refused because they were older than the group's memory of decisions.
scramdb_cluster_dilith_sweeps_resent_totalcounterSweeps sent again with a fresher decision count.
scramdb_cluster_dilith_votes_resent_totalcounterVote requests sent again to groups that had not answered, one per group.
scramdb_cluster_dilith_outcomes_resent_totalcounterTransaction outcomes sent again to groups that had not acknowledged them, one per group.
scramdb_cluster_dilith_aborts_without_record_totalcounterAbort outcomes acknowledged without recording anything at a group holding no record of the transaction.
scramdb_cluster_dilith_decisions_forgotten_totalcounterDecisions forgotten once outside the decision window.
scramdb_cluster_dilith_staged_vote_rounds_totalcounterVote rounds that asked the groups likely to refuse before the others.
scramdb_cluster_dilith_copy_stage_entries_totalcounterCOPY stage entries applied on this node's replicas.
scramdb_cluster_dilith_copy_stages_dropped_totalcounterCOPY stages dropped on this node's replicas without a commit.
scramdb_cluster_dilith_copy_stage_vote_refusals_totalcounterVotes refused because the COPY rows they name were not all staged or conflicted.
scramdb_cluster_dilith_copy_stages_collected_totalcounterCOPY stages a group leader on this node dropped after their owner stopped renewing them.
scramdb_cluster_dilith_copy_stages_refused_full_totalcounterCOPY stage entries a group leader on this node refused because its staging area had no room.
scramdb_cluster_dilith_copy_stages_refused_stale_totalcounterCOPY stage entries a group leader on this node sent back to be checked again because the table moved on since their rows were checked.
scramdb_cluster_dilith_owed_commits_collected_totalcounterCommits whose discharge was lost that a group leader on this node collected.
scramdb_cluster_dilith_table_entriesgaugeEntries each commit protocol table on this node holds, by table.
scramdb_cluster_dilith_owed_commitsgaugeCommits on this node's replicas not yet discharged.
scramdb_cluster_dilith_pending_votesgaugeVotes accepted on this node's replicas and not yet decided.
scramdb_cluster_dilith_recent_verdictsgaugeRecent transaction outcomes this node remembers to answer recoveries.
scramdb_cluster_dilith_discharges_queuedgaugeDischarge notices waiting for their batch to be sent.
scramdb_cluster_dilith_refusing_groupsgaugeGroups that refused this node's votes recently.
scramdb_cluster_dilith_state_bytesgaugeMemory the commit protocol's tables hold on this node, by part: pending_votes (votes accepted and not decided, with their writes and read fences), owed_commits (commits not yet discharged), kept (the decisions and versions the decision and retention windows keep, and COPY stage records), outcomes (outcomes and discharges a group leader holds for its next entry or has in flight), reserved (room reserved for vote appends in flight), waiting (requests parked at a group, and votes waiting for room), scratch (the node's reusable decode and step buffers).
scramdb_cluster_dilith_state_share_bytesgaugeMemory the commit protocol's tables may hold on this node ([cluster.dilith] protocol_state_memory).
scramdb_cluster_dilith_state_peak_bytesgaugeMost memory the commit protocol's tables and reserved room have held on this node.
scramdb_cluster_dilith_votes_waiting_for_memorygaugeVotes waiting on this node for the commit protocol's memory share to have room.
scramdb_cluster_dilith_memory_parks_totalcounterVotes a group leader on this node held until the commit protocol's memory share had room.
scramdb_cluster_dilith_memory_past_share_totalcounterVotes a group leader on this node let past a full memory share so the oldest waiting transaction could go on.

aborts_total{reason="conflict"} and {reason="wait_die"} count contention between transactions on the same rows. A recoveries_started_total that keeps growing on a healthy network means transactions are being left undecided, and is worth a look at the nodes' logs. copy_stages_collected_total rising means COPY statements stop or lose their node while staging; copy_stage_vote_refusals_total rising means a COPY's vote meets concurrent writers of the rows it stages, and the COPY fails with 40001. copy_stages_refused_stale_total counts batches a COPY checked again because its groups moved on while the batch travelled; the COPY goes on. copy_stages_refused_full_total rising means the staging area is too small for the COPYs running (53100). owed_commits_collected_total rising means nodes stop right after answering commits; every such commit is finished by its groups. table_entries shows every table of the commit protocol and what it holds: each one stays bounded by the work in flight.

The commit protocol's tables live in a memory share of their own (state_share_bytes), counted in the process memory budget. A group leader appends a new vote only while the share has room; a vote that finds none waits, oldest transaction first, and is appended as soon as a decision, a discharge or the protocol's windows free room. It is never refused for memory. votes_waiting_for_memory above zero for long, or memory_parks_total climbing, means the share is too small for the transactions in flight: raise [cluster.dilith] protocol_state_memory. memory_past_share_total counts the oldest waiting transaction let through a full share when nothing on the node could free room without it, which is what keeps the cluster moving; its rise says the same.

The commit protocol on this node​

The commit protocol runs as one task on the node's consensus threads, one run at a time, and serves every shard group this node holds a replica of.

MetricTypeMeaning
scramdb_cluster_dilith_run_secondshistogramTime one run of the commit protocol's host took.
scramdb_cluster_dilith_runs_totalcounterRuns of the commit protocol's host.
scramdb_cluster_dilith_run_inputs_totalcounterInputs the commit protocol's host stepped, over all runs.
scramdb_cluster_dilith_queuedgaugeInputs waiting for the commit protocol's host, by kind: control, decide, frame, client.
scramdb_cluster_dilith_step_errors_totalcounterInputs the commit protocol could not step, by input: apply, message, timer, leadership, append_decided, restored, client, hosting.
scramdb_cluster_dilith_reply_slotsgaugeClient calls waiting for an answer from the commit protocol.
scramdb_cluster_dilith_memory_bytesgaugeBytes the commit protocol holds on this node, as last measured.
scramdb_cluster_dilith_appends_in_flightgaugeLog entries the commit protocol proposed whose decision has not come back.
scramdb_cluster_dilith_timersgaugeTimers the commit protocol keeps armed.
scramdb_cluster_dilith_frames_refused_totalcounterFrames the transport refused to send; the protocol sends again.
scramdb_cluster_dilith_appends_lost_totalcounterProposed log entries lost because the group's leader could not take them.
scramdb_cluster_commits_in_doubt_totalcounterCommits whose outcome was not known when the client stopped waiting.
scramdb_cluster_dilith_commits_in_flightgaugeCommits this node's commit protocol drives whose outcome is not decided yet, as last measured.
scramdb_cluster_dilith_commits_abandonedgaugeCommits still in flight here whose client stopped waiting for the outcome, as last measured.
scramdb_cluster_dilith_calls_abandoned_totalcounterCalls to the commit protocol whose caller stopped waiting, forgotten with what waited to answer them.
scramdb_cluster_dilith_stage_rounds_resent_totalcounterRounds of staged rows sent again because their answers did not come in time.
scramdb_cluster_dilith_stage_requests_lost_totalcounterRequests of staged rows asked again at once because their shard group changed leader.
scramdb_cluster_dilith_host_stoppedgauge1 when the commit protocol on this node stopped after a failure.
scramdb_cluster_dilith_hosted_groupsgaugeShard groups the commit protocol hosts on this node.
scramdb_cluster_dilith_groups_hosted_totalcounterShard groups the commit protocol started hosting on this node.
scramdb_cluster_dilith_groups_released_totalcounterShard groups the commit protocol stopped hosting on this node.
scramdb_cluster_dilith_decide_requests_dropped_totalcounterShard group requests dropped because the group is not hosted here or started again.
scramdb_cluster_dilith_node_numbergaugeThis node's number in transaction ids, 0 until the cluster gave it one.
scramdb_cluster_dilith_node_number_claims_totalcounterRequests this node made to the cluster for its node number.
scramdb_cluster_dilith_node_number_waits_totalcounterCommit attempts that waited for this node's node number instead of starting at once.

dilith_host_stopped at 1 means the node can no longer commit transactions and must be restarted; every waiting client was answered with an error. A commits_in_doubt_total that grows means clients stopped waiting before their commit's outcome was known: the commit may still have happened. stage_rounds_resent_total grows when a node sending a COPY's rows, or a transaction's rows too large for one commit message, lost a round to a leader change it could not see or to a node that stopped; one that grows while no leader changed means rounds are taking far longer than the rounds before them.

Committed entries through the commit protocol​

Each shard group hands its committed log entries to the commit protocol, which decides the rows they install, and hands the protocol's own entries to the group's log.

MetricTypeMeaning
scramdb_cluster_dilith_decide_requests_totalcounterBatches of committed entries the data groups handed to the commit protocol.
scramdb_cluster_dilith_decide_entries_totalcounterCommitted entries the data groups handed to the commit protocol.
scramdb_cluster_dilith_decided_rows_totalcounterRows the commit protocol decided for the data groups to install.
scramdb_cluster_dilith_decide_outstandinggaugeRequests of the data groups the commit protocol has not answered yet.
scramdb_cluster_dilith_decide_secondshistogramTime from handing committed entries to the commit protocol to receiving their rows.
scramdb_cluster_dilith_decide_failures_totalcounterData groups stopped because the commit protocol could not decide an entry of theirs.
scramdb_cluster_dilith_leadership_reports_totalcounterLeader changes of the data groups reported to the commit protocol.
scramdb_cluster_dilith_proposal_batches_totalcounterBatches of commit protocol entries handed to a data group's log.
scramdb_cluster_dilith_proposal_batches_refused_totalcounterBatches of commit protocol entries a data group's log could not take.

Applying committed writes​

A committed entry is applied in two steps on the node's apply threads ([cluster.apply] threads), never on the threads that run consensus, replication and client connections. First it is decided: the commit protocol steps it in memory, in log order (the session that committed it already has its answer, given once every group the transaction touched held a durable yes vote). Then its rows are written, and one data log sync per batch makes a whole batch of them durable before the group's applied position moves. Reads wait for that applied position, so a row is never read before it is applied and synced. On a cluster node the consensus threads keep a share of the cores of their own ([cluster.apply] reserve_io_cores), so heavy queries cannot delay them.

MetricTypeMeaning
scramdb_cluster_apply_decide_to_durable_secondshistogramTime from deciding a batch of committed entries to its effects being durable: how far the applied position trails the answers sessions already have.
scramdb_cluster_apply_effect_batches_totalcounterBatches of committed entries whose effects were made durable.
scramdb_cluster_apply_effect_entries_totalcounterCommitted entries whose effects were made durable. Divide by apply_effect_batches_total for the mean batch size.
scramdb_cluster_apply_data_syncs_totalcounterData log syncs the apply path requested for those effects. Divide by apply_effect_batches_total: one per batch (zero for a batch with nothing to write) is the design; more means a write path is syncing on its own.
scramdb_cluster_apply_last_batch_entriesgaugeEntries in the most recent effect batch.
scramdb_cluster_apply_last_batch_data_syncsgaugeData log syncs the most recent effect batch requested.
scramdb_cluster_apply_decided_batches_waitinggaugeDecided batches waiting for their effects. Climbing steadily means the disk or the apply threads cannot keep up; each group's backlog stays bounded by its apply mailbox either way.
scramdb_cluster_apply_stops_totalcounterGroups whose apply stopped on an error. Such a replica is quarantined and rebuilt; see the logs for the reason.
scramdb_cluster_apply_pool_threadsgaugeApply threads running.
scramdb_cluster_apply_pool_queuedgaugeGroups waiting for an apply thread.
scramdb_cluster_apply_effect_slots_reused_totalcounterInstalled rows whose effect was built in a slot an apply thread keeps, its key buffer reused.
scramdb_cluster_apply_effect_slots_fresh_totalcounterInstalled rows whose effect needed a new slot, so its key buffer allocated. Past warm-up it grows only when a batch is larger than any before it on that thread.
scramdb_cluster_apply_effect_images_decoded_totalcounterInstalled rows whose image decode had to allocate memory of its own.
scramdb_cluster_io_runtime_busy_microseconds_totalcounterTime the consensus and connection threads spent running work. Its rate divided by io_runtime_workers is their utilization.
scramdb_cluster_io_runtime_workersgaugeConsensus and connection threads.

Replication flow control​

A leader streams entries to each follower inside a bounded in-flight window (max_inflight_bytes and max_inflight_msgs in [cluster.consensus]), which by default adapts under that ceiling to the bandwidth-delay product measured toward each follower. Each follower's stream is in one of three states: probe (finding where the follower's log ends, one message outstanding), replicate (streaming inside the window) or snapshot (a snapshot is outstanding). The counters and gauges below sum every group this node leads.

MetricTypeMeaning
scramdb_cluster_replication_transitions_totalcounterFollower stream state changes, labelled state (probe, replicate, snapshot) and cause: into probe by leader_start, reject (the follower refused an append or a snapshot), connection_loss or stall (two heartbeat checks with entries in flight and no answer); into replicate by acknowledged or snapshot_installed; into snapshot by behind_compaction.
scramdb_cluster_replication_sends_paused_totalcounterSends held back because a follower's in-flight window was full.
scramdb_cluster_replication_append_messages_totalcounterAppendEntries messages sent to followers, heartbeats included.
scramdb_cluster_replication_append_entries_totalcounterLog entries sent to followers.
scramdb_cluster_replication_append_bytes_totalcounterLog entry bytes sent to followers. Divided by the bytes proposed and by the follower count, this is the copies of each entry the network carried: about one in steady operation.
scramdb_cluster_replication_rejects_totalcounterAppend and snapshot rejects received from followers.
scramdb_cluster_replication_snapshot_resends_totalcounterSnapshots sent again to a follower that had not answered within the resend backoff.
scramdb_cluster_replication_inflight_messagesgaugeUnacknowledged AppendEntries messages to followers, right now.
scramdb_cluster_replication_inflight_bytesgaugeUnacknowledged log entry bytes sent to followers, right now.
scramdb_cluster_replication_window_bytesgaugeThe in-flight window toward a peer, summed over the groups this node leads, labelled peer and reason: ceiling (the configured max_inflight_bytes, or not yet measured), measured (twice the measured bandwidth-delay product), floor (two of the stream's own messages) or pinned (adaptive_inflight_window = false).
scramdb_cluster_replication_round_trip_millisecondsgaugeThe smoothed replication round trip to a peer, from sending an AppendEntries to its acknowledgement (the follower's fsync included), as last measured.
scramdb_cluster_replication_entries_committed_before_leader_sync_totalcounterLog entries committed on followers' synced copies before the leader's own copy reached disk (leader_sends_before_sync in [cluster.consensus]).
scramdb_cluster_replication_leader_unsynced_entriesgaugeLog entries the groups this node leads have sent but not yet synced to their own disk.
scramdb_cluster_replication_lag_entriesgaugeLog entries a peer has not acknowledged, summed over the groups this node leads, labelled peer. A peer whose lag keeps growing is falling behind.
scramdb_cluster_replication_leader_sync_secondshistogramTime from a leader writing new log entries to their sync to disk.
scramdb_cluster_election_timeout_min_millisecondsgaugeLower bound of the election timeout in force (0 until a replication group has started on this node).
scramdb_cluster_election_timeout_max_millisecondsgaugeUpper bound of the election timeout in force.

sends_paused_total climbing together with window_bytes{reason="ceiling"} toward a peer says the ceiling, not the link, is the limit: raise max_inflight_bytes for a long, fast link. Steady growth of transitions_total{cause="stall"} or {cause="reject"} on a healthy network says a follower keeps losing its place and is worth a look at its disk and its connection.

Consensus core buffers​

Consensus reuses its working buffers across steps, so a steady stream of heartbeats, appends and commits allocates nothing for its step output; only a burst above the recent load, or the first load after a quiet spell, needs a fresh allocation, and that spare capacity is handed back again once the burst passes. A received frame is decoded without copying the commands and snapshot data it carries. These figures are for the whole process, summed across it.

MetricTypeMeaning
scramdb_cluster_raft_step_vectors_reused_totalcounterRaft step output vectors served from a kept buffer instead of a new allocation. Counted as each vector is taken.
scramdb_cluster_raft_step_vectors_fresh_totalcounterRaft step output vectors started without a kept buffer, so their first element allocates. Counted as each vector is taken. It stays flat on a steady load and grows when the load keeps more step output in flight than it recently did (a burst, or the first load after a quiet spell).
scramdb_cluster_raft_step_buffer_bytesgaugeBytes the consensus threads keep for Raft step output.
scramdb_cluster_raft_frame_blob_bytes_shared_totalcounterCommand and snapshot bytes decoded from received Raft frames without a copy.
scramdb_cluster_raft_frame_blob_bytes_copied_totalcounterCommand and snapshot bytes copied out of received Raft frames. It stays at zero on a node: every receive path decodes from the frame it was given.
scramdb_cluster_raft_frames_encoded_totalcounterCluster messages encoded into a consensus thread's kept frame buffer.
scramdb_cluster_raft_frame_buffer_growths_totalcounterTimes a consensus thread's kept frame buffer grew to hold a message.
scramdb_cluster_raft_frame_entry_vectors_decoded_totalcounterEntry vectors the Raft frame decoder allocated for received appends, one per append that carries entries.

Heartbeats and quiet groups​

A node sends one heartbeat frame per peer node each heartbeat interval, carrying a line for every group it leads with that peer, instead of one heartbeat per group ([cluster.consensus] coalesced_heartbeats). A frame names its group set by a hash and lists the groups in full only after the set changed. When a group's commit index advances, a read needs confirming or a small append waits, the frame leaves early: at once while the connection to that peer has nothing queued, otherwise after heartbeat_flush_delay. A small append rides inside the frame instead of taking a frame of its own (carry_small_appends), and a follower's reply to an append can ride back inside the answer frame it sends anyway. An idle group (nothing in flight, every member caught up) goes quiet: it stops its heartbeats and its members' election timers, and its peer is refreshed once per election_timeout_min (quiescence). See [cluster.consensus].

MetricTypeMeaning
scramdb_cluster_heartbeat_frames_sent_totalcounterHeartbeat frames this node sent, each covering every group it leads with that peer.
scramdb_cluster_heartbeat_group_lines_sent_totalcounterGroup lines the heartbeat frames of this node carried.
scramdb_cluster_heartbeat_bytes_sent_totalcounterBytes of heartbeat and answer frames this node sent.
scramdb_cluster_heartbeat_full_lists_sent_totalcounterHeartbeat frames that listed every group this node leads with the peer, after that set changed.
scramdb_cluster_heartbeat_quiet_frames_sent_totalcounterHeartbeat frames with no group line, sent to keep a peer's quiet groups quiet.
scramdb_cluster_heartbeat_flushes_totalcounterHeartbeat frames sent early because a commit index advanced, a read needed confirming or a small append waited.
scramdb_cluster_heartbeat_immediate_flushes_totalcounterEarly heartbeat frames sent at once because the connection to the peer had nothing queued.
scramdb_cluster_heartbeat_flush_delay_secondshistogramTime from a commit index advancing, or a read or small append waiting, to the heartbeat frame carrying it.
scramdb_cluster_heartbeat_answers_sent_totalcounterHeartbeat answer frames this node sent to the nodes leading its groups.
scramdb_cluster_heartbeat_frames_received_totalcounterHeartbeat frames this node received.
scramdb_cluster_heartbeat_answers_received_totalcounterHeartbeat answer frames this node received.
scramdb_cluster_heartbeat_round_trip_secondshistogramTime from sending a heartbeat frame to a peer to the peer's answer.
scramdb_cluster_heartbeat_carried_appends_totalcounterSmall appends sent inside heartbeat frames instead of frames of their own.
scramdb_cluster_heartbeat_carried_bytes_totalcounterBytes of the appends sent inside heartbeat frames.
scramdb_cluster_heartbeat_append_replies_totalcounterReplies this node's groups sent to appends of the nodes leading them.
scramdb_cluster_heartbeat_append_reply_frames_totalcounterFrames this node sent for its groups' append replies: a reply sent on its own, or an answer frame carrying replies and no heartbeat answer.
scramdb_cluster_heartbeat_carried_replies_totalcounterAppend replies sent inside heartbeat answer frames instead of frames of their own.
scramdb_cluster_heartbeat_carried_replies_received_totalcounterAppend replies this node received inside heartbeat answer frames.
scramdb_cluster_heartbeat_send_refused_totalcounterHeartbeat and answer frames the transport refused because its queue was full or the peer was not connected.
scramdb_cluster_heartbeat_delivery_refused_totalcounterHeartbeat lines a group could not take because its queue was full or it was gone.
scramdb_cluster_heartbeat_malformed_frames_totalcounterHeartbeat frames dropped because they did not decode.
scramdb_cluster_heartbeat_stale_views_totalcounterHeartbeat frames whose group-set hash did not match the set this node held, which asks for the full list.
scramdb_cluster_heartbeat_quiesce_entered_totalcounterGroups this node leads that went quiet: every member caught up and idle, no heartbeats and no timers.
scramdb_cluster_heartbeat_quiesce_left_totalcounterQuiet groups this node leads that woke up.
scramdb_cluster_heartbeat_woken_by_set_change_totalcounterQuiet groups on this node woken because the node leading them stopped listing them.
scramdb_cluster_heartbeat_woken_by_missed_refresh_totalcounterQuiet groups on this node woken because the node leading them missed two refreshes.
scramdb_cluster_heartbeat_led_groupsgaugeGroups this node leads whose heartbeats its frames carry.
scramdb_cluster_heartbeat_quiescent_groupsgaugeGroups this node leads that are quiet.
scramdb_cluster_heartbeat_peersgaugePeer nodes this node sends heartbeat frames to or receives them from.
scramdb_cluster_heartbeat_interval_millisecondsgaugeThe heartbeat interval in force (heartbeat_interval).
scramdb_cluster_heartbeat_flush_delay_microsecondsgaugeHow long an early heartbeat frame waits to gather more groups, in force (heartbeat_flush_delay).
scramdb_cluster_heartbeat_quiet_refresh_millisecondsgaugeHow often a peer whose groups here are all quiet is refreshed, in force.

group_lines_sent_total divided by frames_sent_total is how many groups one frame carries: the saving over one heartbeat per group. On an idle cluster quiescent_groups approaches led_groups and the heartbeat traffic falls to one small refresh per peer. woken_by_missed_refresh_total rising on a healthy network means a leader's node stalls; send_refused_total, delivery_refused_total and malformed_frames_total stay at zero on a healthy node.

Read-index forwarding​

A replica that is not its group's leader, a follower or a learner, forwards its read-index requests to the leader, which answers them as it answers its own. A forwarded request with no answer within two election timeouts plus the round trip the replica measured to the leader, or whose connection to the leader closed, is refused. An answer that arrives after its request was refused still measures the round trip, so on a slow link the next request waits long enough; forward_late_total rising means a replica's link to its leader is slower than two election timeouts.

MetricTypeMeaning
scramdb_cluster_read_index_forwarded_totalcounterRead-index requests this node forwarded to a group's leader.
scramdb_cluster_read_index_forward_unanswered_totalcounterForwarded read-index requests that ended without the leader's answer.
scramdb_cluster_read_index_forward_late_totalcounterLeader answers to forwarded read-index requests that arrived after the request was given up.
scramdb_cluster_read_index_served_for_peers_totalcounterRead-index requests this node answered as a group's leader for another replica.
scramdb_cluster_read_index_forward_secondshistogramTime from forwarding a read-index request to the leader's answer.

Requests to the metadata and shard groups​

Every request a node sends its metadata group or one of its shard groups has a deadline adapted to the group's election timing. A request the group did not take in time changed nothing and is retried by its caller; a proposal whose answer did not come may still commit, so it is never sent again; a read confirmation is asked again until the deadline. A DDL statement forwarded to the metadata group's leader is bounded the same way. Process-wide, always present, all zero on a healthy cluster:

MetricTypeMeaning
scramdb_group0_reply_timeouts_totalcounterMetadata group proposals whose answer did not come in time; each may still commit.
scramdb_group0_mailbox_timeouts_totalcounterMetadata group requests the group did not take in time; none changed anything.
scramdb_group0_read_retries_totalcounterMetadata group read confirmations asked again after an answer did not come.
scramdb_group0_apply_wait_timeouts_totalcounterWaits for this node to apply a committed metadata group entry that ran out of time.
scramdb_shard_group_reply_timeouts_totalcounterShard group proposals whose answer did not come in time; each may still commit.
scramdb_shard_group_mailbox_timeouts_totalcounterShard group requests the group did not take in time; none changed anything.
scramdb_shard_group_read_retries_totalcounterShard group read points asked again after an answer did not come.
scramdb_shard_group_apply_wait_timeouts_totalcounterWaits for this node to apply a committed shard group entry that ran out of time.
scramdb_ddl_forward_answer_timeouts_totalcounterForwarded metadata group proposals whose answer did not come in time.
scramdb_ddl_forward_in_doubt_totalcounterForwarded metadata group proposals whose leader could not learn their outcome.

reply_timeouts_total or apply_wait_timeouts_total rising means a group is slow to commit or a node is slow to apply; mailbox_timeouts_total rising means the group itself is falling behind on incoming requests.

A metadata group proposal this node's answer did not reach in time is sent again under its first attempt's id, so a leader that already applied it once applies it once, never twice, and a resend past its own deadline is refused rather than applied late:

MetricTypeMeaning
scramdb_group0_proposal_resends_totalcounterMetadata proposals sent again under the id of their first attempt.
scramdb_group0_proposal_duplicates_totalcounterMetadata log entries applied as a repeat of a proposal already applied.
scramdb_group0_proposal_expired_totalcounterMetadata proposals refused because they arrived after their deadline.
scramdb_group0_proposal_window_idsgaugeMetadata proposal ids this node keeps to answer a repeat.

Read positions​

Every read of a cluster table happens at a read position of the commit protocol. A fresh position first has each group's leader confirm the group's latest commit, brings a replica up to it, and waits out any commit still in flight on what is read, so a read never misses a commit already acknowledged to any client, on any node; a reader then serves each group it scans at that position before it reads. For a group this node holds no replica of, the position is asked of the group's leader first, then of its followers, and of its learners last: the leader has nothing to catch up, while a learner may trail its group by any amount, so a learner that has fallen behind does not hold up a statement while a voter of the group can answer. A replica that refuses the position (it does not run the group yet, it is starting, it could not catch up in time; counted on that node by scramdb_cluster_read_position_refusals_total), or fails while still connected, is passed over for the group's next replica. When no replica of such a group is connected to this node (a partition that just healed while its connections form again, a node restarting), the read waits for one to connect (scramdb_cluster_read_position_connect_waits_total). When every replica of the group has refused, as happens for a table created a moment before whose replicas are still starting, the read asks them all again after a pause of a few round trips that grows with each try, and serves the group itself once its own replica of it runs (scramdb_cluster_read_position_reask_waits_total). Either way the read waits for at most its read wait, and then fails with 40001 naming the group. Nothing here writes to a log unless a group is behind the position, where one fence entry brings it up (scramdb_cluster_dilith_fences_appended_total).

MetricTypeMeaning
scramdb_cluster_read_positions_totalcounterFresh read positions chosen by this node.
scramdb_cluster_read_position_secondshistogramTime to choose a fresh read position, from the first barrier to the last answer.
scramdb_cluster_read_position_remote_groups_totalcounterShard groups whose fresh read position another node took for this node (groups this node holds no replica of).
scramdb_cluster_read_position_connect_waits_totalcounterFresh read positions that waited for a replica of a shard group to connect.
scramdb_cluster_read_position_reask_waits_totalcounterFresh read positions that asked a shard group again after every replica refused it.
scramdb_cluster_read_position_refusals_totalcounterFresh read positions refused here because this node's replica could not serve them.
scramdb_cluster_read_position_raised_totalcounterTransactions whose read position rose to reach a table they had not read before.
scramdb_cluster_read_barrier_retries_totalcounterRead barriers asked again because a shard group had no leader to confirm them.
scramdb_cluster_read_serves_totalcounterShard groups served at a read position on this node.
scramdb_cluster_read_serve_secondshistogramTime a shard group took to reach a read position on this node.
scramdb_cluster_read_failures_totalcounterReads refused because a shard group could not be confirmed or served in time.
scramdb_cluster_read_registered_ranges_totalcounterTable reads registered with a transaction's commit for validation.
scramdb_cluster_read_skip_serves_totalcounterShard groups served for SKIP LOCKED or NOWAIT reads without waiting for a commit.
scramdb_cluster_read_skip_waits_totalcounterSKIP LOCKED or NOWAIT reads of a shard group that waited for a commit in flight.
scramdb_cluster_read_skip_fence_waits_totalcounterWaits of SKIP LOCKED or NOWAIT reads for a shard group to be fenced at their cut.
scramdb_cluster_read_skipped_keys_totalcounterKeys a SKIP LOCKED or NOWAIT read found written by a commit in flight.
scramdb_cluster_read_skipped_rows_totalcounterRows SKIP LOCKED left out because a commit in flight writes them.
scramdb_cluster_read_nowait_refusals_totalcounterNOWAIT statements refused because a commit in flight writes a row they select.
scramdb_cluster_read_join_spills_totalcounterReads over more shard groups than fit inline, whose answers spilled to the heap: reads that touch more than eight groups at once.
scramdb_cluster_read_horizon_serves_totalcounterShard group watermarks taken here for a transaction snapshot, waiting for no commit.
scramdb_cluster_read_unreached_groups_totalcounterShard groups a transaction snapshot could reach on no replica and left out.
scramdb_cluster_read_fence_waits_totalcounterWaits for a shard group to be fenced at a read position, waiting for no commit.
scramdb_cluster_read_range_serves_totalcounterKey ranges served at a read position that waited only for commits writing them.

A read that finds no leader for a group keeps asking until its wait runs out (one election timeout plus the time a group may hold requests without a leader plus the time an undecided transaction may block it, about ten seconds with the defaults) and is then refused with SQLSTATE 40001; a replica that cannot catch up is refused with 57P03. Growth of read_failures_total together with read_barrier_retries_total means a group is without a leader. read_position_raised_total counts transactions that read a table only after an earlier statement had fixed their position: such a transaction's earlier reads are validated at their own position when it commits. Under READ COMMITTED it also counts statements whose position rose because a table they reached later, such as a foreign key's parent, was fresher, so a row committed before the statement began is never missed.

SKIP LOCKED and NOWAIT read the tables they lock without waiting for any commit in flight: the keys such commits write come back with the rows (read_skipped_keys_total), and the statement leaves those rows out (read_skipped_rows_total) or refuses with 55P03 (read_nowait_refusals_total). A shard group where a commit in flight writes a table the statement reads without locking it, such as a join's other side, is read the ordinary way, waiting for that commit, and counts in read_skip_waits_total.

A REPEATABLE READ or SERIALIZABLE transaction takes its snapshot from every shard group's confirmed watermark without waiting for any commit in flight (read_horizon_serves_total), then fences the groups this node holds at it (read_fence_waits_total counts the waits for a fence on its way). A read of a few keys at that snapshot waits only for a commit in flight on those keys (read_range_serves_total). A group the snapshot can reach on no replica, such as one cut off by a partition, is left out (read_unreached_groups_total): the transaction still reads every other table, and its reads are checked at COMMIT, which fails with 40001 if one of them missed a commit. A group whose replicas the snapshot reaches but which all refuse it, such as the groups of a table created a moment before whose replicas are still starting, is not left out: it is asked again at the horizon retry interval (horizon_retry, 90 ms by default) until a replica answers, for at most the read wait.

Adaptive routing​

MetricTypeMeaning
scramdb_cluster_routes_totalcounter/labelRouting decisions, labelled route: one per statement, and one per scatter for a statement that scatters more than once (each outer row of a k-NN LATERAL join).

route takes exactly one of five values: oltp, point, local_replica, forward_replica, scatter. That's the whole label space, fixed by construction, the keys are the five route kinds the planner can choose, never a node id or a group id, so this metric's cardinality can never grow with the size of the cluster. A route kind the counter doesn't recognize is dropped rather than bucketed into a sixth catch-all series, so a routing path that forgot to name itself shows up as a silent gap here, not a mislabeled entry.

Where fragments run​

A distributed read runs each bucket's fragment on one replica that holds the bucket: a learner before a follower before the shard group's leader, in this node's region first, spread evenly (see Where fragments run). [cluster] fragment_any_replica = false keeps fragments on the buckets' owners.

MetricTypeMeaning
scramdb_cluster_exchange_placed_learner_buckets_totalcounterBuckets of distributed reads this node placed on a learner replica.
scramdb_cluster_exchange_placed_follower_buckets_totalcounterBuckets of distributed reads this node placed on a follower replica.
scramdb_cluster_exchange_placed_leader_buckets_totalcounterBuckets of distributed reads this node placed on a shard group's leader.
scramdb_cluster_exchange_placement_skips_totalcounterReplicas placement passed over because they could not serve the read position.
scramdb_cluster_exchange_fragment_redispatches_totalcounterFragments run on another replica after one could not serve them.
scramdb_cluster_exchange_fragments_served_learner_totalcounterFragments this node served from its learner store.
scramdb_cluster_exchange_round_replacements_totalcounterShuffle and two-phase read rounds run again on other replicas after a node could not serve its buckets.
scramdb_cluster_exchange_fragments_refused_totalcounterFragments this node refused because its replica could not serve the read position.
scramdb_cluster_exchange_shipped_tables_totalcounterTables this node sent with a query's fragments because no replica held them whole.

placed_leader_buckets_total growing against the learner and follower counts means the learners and followers cannot serve the reads (behind their group, or waiting for a full copy of its rows), so analytical scans land on the leaders; fragments_refused_total that stays high on one node says that node's replicas lag.

Distributed shuffle joins​

How a shuffle join across nodes adapted while it ran ([cluster.distributed_join] in the config file).

MetricTypeMeaning
scramdb_cluster_shuffle_consumer_fragments_totalcounterConsumer fragments dispatched for a materialized shuffle join, one per coalesce group. Equals the participant count with coalescing off; lower with it on. Zero on a streaming query.
scramdb_cluster_shuffle_aqe_coalesce_groups_totalcounterCoalesce groups actually acted on: each merged more than one small adjacent partition into a single consumer fragment.
scramdb_cluster_shuffle_aqe_skew_split_flagged_totalcounterPartitions flagged skewed (over 5x the median and over 256MB) but not acted on: skew-split off, or a single-node participant set with nothing to split across. Logged and metered even when nothing follows.
scramdb_cluster_shuffle_aqe_skew_split_acted_totalcounterSub-consumers dispatched for a skew-split partition that was acted on: the larger side split into disjoint row ranges, the smaller side replicated. Summed across every acted partition.
scramdb_cluster_shuffle_aqe_broadcast_demote_flagged_totalcounterBroadcast sides flagged over budget, a broadcast-demote candidate. Logged and metered, not acted on yet.
scramdb_cluster_shuffle_straggler_backups_totalcounterSpeculative backup consumer fragments dispatched because a primary consumer didn't return within the straggler-backup delay, on a different live node reading the same frozen spools.

Node operations and cluster TRUNCATE​

The whole-cluster vector statistics views and a cluster TRUNCATE reach every other node through the same exchange the distributed queries use.

MetricTypeMeaning
scramdb_cluster_exchange_node_ops_sent_total / scramdb_cluster_exchange_node_ops_failed_totalcounterNode operations this node asked a peer to run, and those a peer could not run or never answered.
scramdb_cluster_exchange_node_ops_served_total / scramdb_cluster_exchange_node_ops_serve_failed_totalcounterNode operations this node ran for a peer, and those it could not.
scramdb_cluster_exchange_view_gathers_total / scramdb_cluster_exchange_view_nodes_unreachable_totalcounterWhole-cluster statistics view reads this node gathered, and nodes such a read reported as unreachable.
scramdb_cluster_exchange_view_gather_secondshistogramTime to gather one whole-cluster statistics view read.
scramdb_cluster_exchange_truncates_total / scramdb_cluster_exchange_truncate_failures_totalcounterCluster TRUNCATE statements this node coordinated, and those that did not complete on every node.
scramdb_cluster_exchange_truncate_secondshistogramTime for one cluster TRUNCATE to reach every node holding the table.
scramdb_cluster_exchange_producer_failures_propagated_total / scramdb_cluster_exchange_producer_failure_undelivered_totalcounterFailed shuffle producers that handed their error to every consumer at once, and consumers they could not reach (those fall back to their own receive deadline).

Cluster branches and table commands​

A CREATE DATABASE ... CLONE or a scram.branch_at on a cluster, and the table commands that carry a branch, a CLONE's fork hold and a cluster TRUNCATE through each shard group's log.

MetricTypeMeaning
scramdb_cluster_branches_created_totalcounterBranch databases this node created on the cluster.
scramdb_cluster_branch_failures_totalcounterBranch databases this node could not create on the cluster.
scramdb_cluster_branch_tables_forked_totalcounterTables forked into a branch database on this node's stores.
scramdb_cluster_branch_pending_rows_totalcounterRows a shard group gave a branch after this node's store forked it, before the group's fork entry.
scramdb_cluster_branch_shared_stages_totalcounterCOPY segments shared into a branch as they attached after this node's store forked it.
scramdb_cluster_branch_holds_taken_totalcounterFork holds taken on this node's stores, each keeping row versions until its branch is forked.
scramdb_cluster_branch_holds_released_totalcounterFork holds released on this node's stores.
scramdb_cluster_branch_clones_abandoned_totalcounterCLONE statements this node abandoned, their fork holds released on every node.
scramdb_cluster_branch_archive_builds_totalcounterBranches at a past time this node built from the cluster archive.
scramdb_cluster_branch_archive_rows_totalcounterRows this node read from the cluster archive into a branch at a past time.
scramdb_cluster_dropped_database_rows_skipped_totalcounterRows of a dropped database a shard group skipped on this node.
scramdb_cluster_branch_fork_recordsgaugeBranch fork entries the shard groups on this node still record one by one. A record is folded away once the store holding its group settled the branch (made it ready, or dropped it), so this stays at the branches in progress.
scramdb_cluster_branch_fork_records_pruned_totalcounterBranch fork records folded into a group's floor once their branch was settled.
scramdb_cluster_dropped_databases_recordedgaugeDropped databases this node's stores record one by one; each is released once its drop is decided.
scramdb_cluster_table_commands_proposed_totalcounterTable commands this node placed in a shard group's log.
scramdb_cluster_table_command_failures_totalcounterTable commands this node could not place in a shard group's log before the deadline.
scramdb_cluster_table_command_redirects_totalcounterTable command proposals a node turned away because its replica did not lead the group.

Table definition changes​

An ALTER TABLE on a cluster: the commits it holds back on each node, the fence it places in each of a distributed table's shard groups, and the rows each node installs across a change.

MetricTypeMeaning
scramdb_cluster_schema_changes_pendinggaugeTables whose definition change is under way on this node's stores.
scramdb_cluster_schema_commits_refused_totalcounterCommits refused because a table they wrote changed its definition while they ran.
scramdb_cluster_schema_commit_drains_totalcounterWaits for the commits under way before a table's definition change.
scramdb_cluster_schema_commit_drain_timeouts_totalcounterWaits for the commits under way that ended at their deadline.
scramdb_cluster_schema_fences_applied_totalcounterTable definition change fences this node's stores applied from shard group logs.
scramdb_cluster_schema_fence_waits_totalcounterFences that waited for this node's catalog to take the change first.
scramdb_cluster_renamed_table_rows_totalcounterRows of a renamed distributed table installed under its new name.
scramdb_cluster_widened_rows_totalcounterRows written before a column was added, widened to the table's shape as installed.

A DROP TABLE of a distributed table holds back and drains the table's commits and fences its shard groups the same way before the table goes, so it moves schema_commit_drains_total and schema_fences_applied_total too.

UPDATE and DELETE on a partially held table​

An UPDATE or DELETE on a node that holds only some buckets of a distributed table reads its rows from every owner (see UPDATE and DELETE across nodes). A node that holds every bucket of the table runs the single-node path and moves none of these.

MetricTypeMeaning
scramdb_cluster_dml_statements_distributed_totalcounterUPDATE and DELETE statements that read their rows from every owner.
scramdb_cluster_dml_candidates_remote_total / scramdb_cluster_dml_candidates_local_totalcounterRows such statements received from other nodes, and read from this node's own buckets.
scramdb_cluster_dml_candidate_batches_totalcounterRow batches such statements processed.
scramdb_cluster_dml_candidate_fetch_secondshistogramTime a statement waited for its next batch of rows.
scramdb_cluster_dml_candidate_spilled_bytes_totalcounterRow bytes set aside on disk because the statement's memory budget was spoken for.
scramdb_cluster_dml_owners_pruned_totalcounterOwners a statement skipped because its WHERE names one bucket.
scramdb_cluster_dml_locking_selects_totalcounterLocking SELECT statements (FOR UPDATE, FOR NO KEY UPDATE, FOR SHARE, FOR KEY SHARE) that read a distributed table's rows without locking them.
scramdb_cluster_dml_locking_read_rows_totalcounterRows such statements took and registered one by one for their commit to check (a SERIALIZABLE statement without SKIP LOCKED has its whole scan checked instead and adds nothing; rows SKIP LOCKED left out are not counted).
scramdb_cluster_dml_own_writes_reads_totalcounterQueries that read a partially held table together with their transaction's own uncommitted changes to it.
scramdb_cluster_dml_candidate_readers_replaced_totalcounterReaders of an UPDATE, DELETE or uniqueness check that could not serve their buckets and were read again on another replica.

A growing candidate_spilled_bytes_total means large UPDATE or DELETE statements are running close to the execution memory budget.

The row-lock family these statements used to report (scramdb_cluster_dml_row_lock_batches_total, scramdb_cluster_dml_row_locks_total, scramdb_cluster_dml_row_lock_waits_total, scramdb_cluster_dml_row_lock_wait_seconds, scramdb_cluster_dml_row_lock_conflicts_total, scramdb_cluster_dml_row_locks_released_total) and scramdb_cluster_dml_statement_restarts_total are gone: rows of a distributed table are no longer locked, so no statement waits, restarts or releases a lock (see Row locks on distributed tables). Writers that reach the same rows now show up as 40001 failures at commit: watch scramdb_cluster_dilith_aborts_total and your clients' retry counts instead, and remove any alert or dashboard panel that reads the removed names.

Distributed statements​

MetricTypeMeaning
scramdb_cluster_read_stability_retries_totalcounterReads re-run because a replica applied writes while they scanned.
scramdb_cluster_read_stability_failures_totalcounterReads that failed because replicas kept applying writes through every re-run. The statement fails with 40001 and can be retried.
scramdb_cluster_serialization_retries_totalcounterAutocommit statements sent again, and COPY votes voted again, after a serialization failure.
scramdb_cluster_placement_waits_totalcounterStatements that waited for this node to learn where a table they name is placed, right after it restarted or right after the table was created.
scramdb_cluster_rows_moved_totalcounterRows an UPDATE moved to a new primary key or shard group.
scramdb_cluster_unique_probes_totalcounterUniqueness checks of new keys sent to the nodes that hold their buckets.
scramdb_cluster_unique_violations_totalcounterNew keys refused because a row on another node already holds them.
scramdb_cluster_inserts_positioned_by_key_totalcounterInserted batches whose keys were checked in their own shard groups only, not in every group of the table.
scramdb_cluster_upsert_remote_conflicts_totalcounterINSERT ... ON CONFLICT rows whose conflicting row is stored on another node.
scramdb_cluster_on_conflict_keys_overtaken_totalcounterINSERT ... ON CONFLICT statements refused because another transaction committed one of their conflict keys after the statement judged it free.
scramdb_cluster_own_writes_spilled_bytes_totalcounterBytes of a transaction's view of a table a query set aside on disk for want of memory.
scramdb_cluster_general_statements_totalcounterStatements answered on their coordinator over rows streamed from every shard group.
scramdb_cluster_general_fast_path_declines_totalcounterStatements a faster distributed path handed to the general path.
scramdb_cluster_general_rows_streamed_totalcounterRows streamed to a coordinator for the statements it answered over them.
scramdb_cluster_general_spilled_bytes_totalcounterBytes of streamed rows a coordinator set aside on disk for want of memory.
scramdb_cluster_general_readers_refused_totalcounterReaders of a streamed relation whose replica refused the read position. Their buckets were read on another replica.
scramdb_cluster_general_readers_unreached_totalcounterReaders of a streamed relation that were not reached or stopped sending. Their buckets were read on another replica.
scramdb_cluster_general_readers_failed_totalcounterReaders of a streamed relation whose read failed, failing the statement.
scramdb_cluster_general_reads_exhausted_totalcounterStreamed reads that failed because no replica was left to read a bucket. The statement fails with 40001 and can be retried.

Prepared transactions​

PREPARE TRANSACTION, COMMIT PREPARED and ROLLBACK PREPARED on a cluster's distributed tables (see Two-phase commit), counted on the node that prepared the transaction.

MetricTypeMeaning
scramdb_cluster_two_phase_prepared_totalcounterTransactions prepared on this node.
scramdb_cluster_two_phase_prepare_failures_totalcounterPREPARE TRANSACTION statements that failed and rolled their transaction back.
scramdb_cluster_two_phase_committed_totalcounterPrepared transactions committed on this node.
scramdb_cluster_two_phase_commit_conflicts_totalcounterCOMMIT PREPARED statements refused with 40001 because a later commit overtook the transaction.
scramdb_cluster_two_phase_commit_unknown_totalcounterCOMMIT PREPARED statements whose outcome was not known when they answered (08007).
scramdb_cluster_two_phase_commit_failures_totalcounterCOMMIT PREPARED statements that failed for another reason, such as an unknown transaction name.
scramdb_cluster_two_phase_placement_waits_totalcounterCOMMIT PREPARED statements that waited, right after this node restarted, for it to learn where their tables are placed.
scramdb_cluster_two_phase_rolled_back_totalcounterPrepared transactions rolled back on this node.
scramdb_cluster_dilith_prepared_attempts_totalcounterCommit attempts PREPARE TRANSACTION kept for a later COMMIT PREPARED.
scramdb_cluster_dilith_prepared_attempts_resumed_totalcounterKept attempts COMMIT PREPARED took up to commit.

A growing commit_conflicts_total means prepared transactions wait long enough for other writers to change what they read; commit them sooner, or retry them from the start.

Columnar mirror​

A node that keeps a columnar copy of the rows it replicates (a hybrid voter, columnar_replica = true, or a learner) copies committed entries into it on its own mirror threads.

MetricTypeMeaning
scramdb_cluster_columnar_mirror_threadsgaugeThreads that copy committed entries into this node's columnar store.
scramdb_cluster_columnar_mirror_queuedgaugeColumnar mirror lanes waiting for a mirror thread.
scramdb_cluster_columnar_mirror_panics_totalcounterColumnar mirror lanes stopped by a panic. Zero on a healthy node; the log names the cause.

Learner-read freshness​

MetricTypeMeaning
scramdb_cluster_learner_freshness_secondshistogramHow long ago each queried group's read floor last advanced, observed once per stale or strong learner_read read. The floor is the commit protocol's closed position at the replica: every commit at or below it is installed there and none can land there later. The OLTP path (learner_read off) never touches this histogram. Same buckets as the handshake histogram, 1ms to 5s plus +Inf, a natural fit for replication lag.

This is the same lag SHOW FRESHNESS reports per data group in Clustering, captured here as a scrapeable distribution instead of a point-in-time value.

Faults, retries and reachability​

How a node sees its peers and its consensus groups, and every retry, timeout and re-dispatch of the cluster's network mechanisms. Always emitted in cluster mode. On a healthy cluster the failure and retry counters stay at zero, and the gauges show the peers and members that are up. A fault that repeats is logged once, as a warning when it begins, and counted every time; a lost peer that is heard again is logged once at info.

MetricTypeMeaning
scramdb_cluster_peers_reachablegaugePeers with an open cluster connection.
scramdb_cluster_peer_reachablegaugeWhether a peer connected or lost now has an open cluster connection (1 or 0), labelled peer.
scramdb_cluster_peer_round_trip_microsecondsgaugeRound trip of the newest answered liveness ping to a peer, labelled peer. Absent until a ping is answered.
scramdb_cluster_peer_silence_millisecondsgaugeTime since a peer's cluster connection last received a byte, labelled peer.
scramdb_cluster_peers_lostgaugePeers that were connected and have not been heard from since their connection was lost.
scramdb_cluster_peer_losses_totalcounterTimes a connected peer became unreachable. A partition is one loss per peer, however often its connections close.
scramdb_cluster_dial_failures_totalcounterConnection attempts to peers that failed, labelled cause.
scramdb_cluster_inbound_handshake_failures_totalcounterConnections from peers that failed or were refused during the handshake, labelled cause.
scramdb_cluster_accept_errors_totalcounterTimes a cluster listener failed to accept a connection.
scramdb_cluster_groups_without_leadergaugeConsensus groups on this node that know no leader.
scramdb_cluster_group_leader_losses_totalcounterTimes a consensus group on this node lost its leader.
scramdb_cluster_stalled_groupsgaugeConsensus groups on this node that made no progress for as long as a request waits for its answer (fifteen election timeouts, at least 5 s).
scramdb_cluster_group_stalls_totalcounterTimes a consensus group on this node stalled.
scramdb_cluster_retries_totalcounterTimes a cluster mechanism tried again after a failure, labelled mechanism.
scramdb_cluster_retries_exhausted_totalcounterTimes a cluster mechanism gave up after its retries or its deadline ran out, labelled mechanism.
scramdb_cluster_timeouts_totalcounterWaits of a cluster mechanism that reached their deadline, labelled mechanism.
scramdb_cluster_redispatches_totalcounterWork placed again on another node after its first node failed, labelled mechanism.
scramdb_cluster_retry_rate_per_secondgaugeRetries a second across every mechanism over the last ten seconds. A sustained rise is a retry storm.
scramdb_group0_reply_timeouts_total / scramdb_shard_group_reply_timeouts_totalcounterMetadata group and shard group proposals whose answer did not come in time. Each may still commit.
scramdb_group0_mailbox_timeouts_total / scramdb_shard_group_mailbox_timeouts_totalcounterRequests the group did not take in time. None changed anything.
scramdb_group0_read_retries_total / scramdb_shard_group_read_retries_totalcounterRead confirmations asked again after an answer did not come.
scramdb_group0_apply_wait_timeouts_total / scramdb_shard_group_apply_wait_timeouts_totalcounterWaits for this node to apply a committed entry that ran out of time.
scramdb_cluster_swim_members_alive / _suspect / _deadgaugeCluster members this node's failure detector sees alive, suspects, and has declared failed.
scramdb_cluster_swim_local_healthgaugeThis node's failure detector health score. Above zero, the node itself is missing probe answers.
scramdb_cluster_swim_member_up_total / _suspect_total / _down_totalcounterTimes a member came up, became suspect, or was declared failed in this node's failure detector.
scramdb_cluster_swim_malformed_messages_totalcounterFailure detector messages from peers that could not be read and were dropped.
scramdb_cluster_swim_send_failures_totalcounterFailure detector messages this node could not hand to a peer's connection.
scramdb_cluster_swim_joins_dropped_totalcounterMember joins dropped because the failure detector's join queue was full.
scramdb_cluster_exchange_partial_chunksgaugeExchange streams holding part of a chunk in reassembly on this node.
scramdb_cluster_exchange_malformed_frames_total / scramdb_cluster_exchange_orphan_frames_totalcounterExchange frames dropped because they could not be read, or because no query on this node was waiting for their stream.
scramdb_cluster_exchange_out_of_order_total / _undecodable_chunks_total / _reassembly_overflows_totalcounterExchange streams failed because a frame arrived out of order, a chunk could not be decoded, or a chunk grew past the reassembly limit.
scramdb_cluster_exchange_producer_failures_received_totalcounterExchange streams whose producer reported a failure to this node.
scramdb_cluster_stream_frames_malformed_totalcounterFrames of a cluster service from peers that could not be read and were dropped.
scramdb_cluster_point_forwards_totalcounterPoint reads this node forwarded to their group's leader.
scramdb_cluster_point_forward_send_failures_total / _remote_errors_total / _reply_send_failures_totalcounterForwarded point reads that could not be sent, that the leader answered with an error, and answers to other nodes' forwarded reads that could not be sent back. A forwarded read that could not be sent or got no answer is answered on the general path instead; an error the leader reported is the statement's answer.

The cause of a failed connection is one of connect_timeout, refused, reset, unreachable, unresolved, handshake_timeout, rejected, version (the two nodes share no cluster protocol version: upgrade the older one), identity, tls, protocol, closed and other.

The mechanism is one of the following, and each family lists only the mechanisms that can record it, so no sample reads as a measured zero:

  • dial and lane_dial: connecting to a peer;
  • point_forward, fragment and fetch: a distributed statement's requests;
  • install: fetching a state snapshot;
  • drain and autojoin: membership changes;
  • metadata_group and shard_group: requests to a consensus group;
  • commit_vote and commit_outcome: the commit protocol's resends.

A fragment timeout is a reader that missed its deadline. Its buckets are read on another replica, which is a fragment re-dispatch.

Alerting on cluster metrics​

Three conditions worth paging on. Each maps to a documented design decision rather than a guessed threshold:

groups:
- name: scramdb-cluster
rules:
- alert: ScramDBGroupsQuarantined
expr: scramdb_cluster_quarantined_groups > 0
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $value }} raft group(s) quarantined"
description: >
A group's storage was quarantined at boot for corruption and has
not finished its membership-replace rejoin. This gauge is
recomputed from the real gated set every reconcile round, not
decremented, so restarting the node will not clear a stale
reading; only a completed rejoin does.

- alert: ScramDBLogCompactionStalled
expr: |
rate(scramdb_cluster_log_compactions_total[15m]) == 0
and rate(scramdb_cluster_group_commit_ops_total[15m]) > 0
for: 30m
labels:
severity: warning
annotations:
summary: "Raft log compaction has not run on a busy node in 30 minutes"
description: >
The node is committing but not compacting. Either its state
snapshots keep failing (check
scramdb_cluster_state_snapshot_failures_total and
scramdb_cluster_state_snapshots_abandoned_total: a full disk or a
memory_bytes budget too small), or a replica is behind but still
inside the lag limit. A replica lagging past the limit stops
holding the log and is repaired from a snapshot when it returns.

- alert: ScramDBPartitionDropsDetected
expr: increase(scramdb_cluster_partition_drops_total[5m]) > 0
labels:
severity: critical
annotations:
summary: "Cluster transport frames are being silently dropped"
description: >
This counter only moves when a network partition is injected
(chaos testing) or something outside the engine is dropping
cluster transport frames. It is zero in every normal deployment,
so any increase is worth paging on.

GIN index metrics​

Process-wide totals over every GIN index (USING gin) on the node, always present, on the same /metrics body. Per-index size is not labelled here; EXPLAIN shows which index a query probes.

MetricTypeMeaning
scramdb_gin_probes_totalcounterGIN index probes run.
scramdb_gin_candidates_totalcounterCandidate rows the posting lists offered.
scramdb_gin_recheck_rejected_totalcounterCandidates the recheck of the query's filter rejected.
scramdb_gin_rows_returned_totalcounterRows GIN probes returned after their recheck.
scramdb_gin_snapshot_scans_totalcounterGIN-planned reads inside a REPEATABLE READ or SERIALIZABLE transaction, answered by scanning the table as of the snapshot.
scramdb_gin_entries_inserted_total / scramdb_gin_entries_deleted_totalcounterIndex entries added / removed by writes.
scramdb_gin_probe_secondshistogramTime one GIN probe took, posting lists to returned rows.

A high recheck_rejected_total against candidates_total means the index narrows poorly for your filters; a rising snapshot_scans_total means snapshot transactions are reading whole tables where a READ COMMITTED read would probe the index.

Index builds and compaction​

An index build and a compaction merge never work on one table at the same time, as in PostgreSQL, where CREATE INDEX and VACUUM take conflicting locks. Background compaction yields: a merge running when a build starts stops, and the table is merged on a later pass. An explicit VACUUM waits for the builds on its table to finish, and a build waits for a running VACUUM, for at most lock_timeout when it is set (then it fails with SQLSTATE 55P03). Process-wide, always present:

MetricTypeMeaning
scramdb_compaction_merges_yielded_to_index_builds_totalcounterBackground merges stopped or skipped because an index build held their table.
scramdb_index_builds_waited_for_merges_totalcounterIndex builds that waited for a merge running on their table to end.
scramdb_index_build_merge_wait_timeouts_totalcounterIndex builds that stopped waiting for a merge on their table at their lock timeout.
scramdb_vacuums_waited_for_index_builds_totalcounterVACUUM runs that waited for the index builds holding their table to end.

A separate fence orders a compaction's row relocation against a concurrent write that resolves a row's physical position by table (DELETE, UPDATE, ON CONFLICT DO UPDATE/DO SELECT, a replica's apply): the write holds the fence until its commit is durable, and a compaction raises it before it swaps its merged segments in, so a swap never relocates a row out from under a write already under way. A statement that reads a table's row identities (a ctid column, a TABLESAMPLE draw) holds the same fence until it ends, or inside a transaction block until the block ends, so each row keeps one ctid for it, as PostgreSQL's VACUUM FULL waits for every transaction that read the table. Process-wide, always present:

MetricTypeMeaning
scramdb_writes_waited_for_compaction_totalcounterStatements that waited for a compaction to finish moving their table's rows.
scramdb_write_compaction_wait_timeouts_totalcounterStatements that stopped waiting for a compaction at their lock timeout.
scramdb_compaction_guarded_swaps_totalcounterCompaction swaps made while the statements using their table's rows were held back.
scramdb_compaction_swaps_abandoned_totalcounterBackground compaction batches dropped because the statements holding their table's rows did not finish in time.
scramdb_vacuum_lock_timeouts_totalcounterVACUUM FULL runs stopped at their lock timeout while statements held the table's rows.
scramdb_compaction_late_deletes_carried_totalcounterRows deleted while a compaction ran that it carried to their new places.
scramdb_compaction_writer_drain_secondshistogramTime a compaction waited for the statements holding its table's rows to finish.
scramdb_compaction_writer_pause_secondshistogramTime a compaction held back the statements using its table's rows while it swapped.
scramdb_write_compaction_wait_secondshistogramTime a statement waited for a compaction before it could read or write its rows.
scramdb_compaction_writer_tablesgaugeTables with writer counters allocated.

swaps_abandoned_total or vacuum_lock_timeouts_total rising means the statements holding a table's rows (its writers, and its ctid or TABLESAMPLE readers) never drain within the wait bound: background compaction backs off quietly, VACUUM FULL fails loudly with 55P03.

Query worker pool​

The engine's core-pinned query worker pool ([execution] workers), process-wide, always present:

MetricTypeMeaning
scramdb_query_worker_countgaugePinned workers the engine's query worker pool runs.
scramdb_query_worker_queue_depthgaugeOutstanding tasks queued across the query worker pool's pinned workers, summed.
scramdb_query_worker_cpu_quota_cores_permillegaugeThis engine's claimed cores per thousand of the machine's full, unconstrained cpuset: 1000 unconstrained, lower under a CPU limit (a container --cpus, a Kubernetes CPU request).

worker_queue_depth climbing while worker_count holds steady means query work is arriving faster than the pool drains it; cpu_quota_cores_permille well below 1000 on a box that looks otherwise idle means the CFS quota, not the workload, is the ceiling on worker_count.

Stand-in query workers​

A fixed-pool query worker whose core is parked on a lock wait hands its core to a stand-in worker for the wait's duration, so a lock contended by one statement does not idle every query on that core. Process-wide, always present:

MetricTypeMeaning
scramdb_query_standinsgaugeStand-in query worker threads alive right now, running or idle-cached.
scramdb_query_standin_starts_totalcounterStand-in query worker threads started.

Vector index metrics​

Every anode/manode vector index family lives on the scramdb_vector_ prefix, appended to the same /metrics body as everything above. One sample feeds both the Prometheus families below and the system views in the next section, so a number never disagrees between them. A family the sample has not measured (a recall job that never ran, a per-index rate with no traffic yet) is omitted from the scrape entirely rather than rendered as a fabricated 0, the same rule the WAL archive metrics above follow. Cardinality is bounded by the number of indexes (and cluster parts) open on the node, reported by scramdb_vector_series so you can see the cost.

Node totals (no index label, always present):

MetricTypeMeaning
scramdb_vector_memory_budget_total_bytesgaugeThe memory ceiling configured for all vector indexes on this node together.
scramdb_vector_memory_reserved_bytesgaugeMemory currently reserved by vector index arenas against the total ceiling.
scramdb_vector_memory_in_use_bytesgaugeMemory vector indexes use right now: their arenas, row maps and write buffers together. The write buffers are charged to execution memory, not to the vector ceiling. A build's working memory (the model training sample, the convert buffers) is not counted here: it is charged to storage.memory.maintenance_bytes, like every index build.
scramdb_vector_memory_engine_bytesgaugeMemory the engine holds for vector index ledgers, bitmaps and buffers.
scramdb_vector_memory_budget_refusals_totalcounterIndex builds, rebuilds and arena regrows refused because the total ceiling would be exceeded.
scramdb_vector_indexesgaugeVector indexes open on this node.
scramdb_vector_regrows_total / scramdb_vector_regrows_runningcounter / gaugeArena regrows completed / running on this node.
scramdb_vector_scans_total{tier}counterIndexed vector searches served, labelled exact, traversal, or postfilter.
scramdb_vector_scan_secondshistogramTime an indexed vector search took end to end.
scramdb_vector_rerank_rows_total / scramdb_vector_rerank_secondscounter / histogramCandidate rows re-ranked exactly after an index search, and how long it took.
scramdb_vector_widening_rounds_totalcounterExtra index rounds taken when a filtered search had fewer than the requested rows.
scramdb_vector_underfilled_totalcounterFiltered searches that returned fewer rows than requested after the scan limit.
scramdb_vector_knn_join_batches_total / scramdb_vector_knn_join_queries_totalcounterBatches / individual query vectors sent to an index by a nearest-neighbour join.
scramdb_vector_threshold_scans_totalcounterDistance threshold searches served by an index.
scramdb_vector_exact_scans_total / scramdb_vector_exact_scan_secondscounter / histogramNearest-neighbour statements served by the exact scan without an index, and how long it took.
scramdb_vector_distributed_scans_total / scramdb_vector_distributed_merge_rows_totalcounterNearest-neighbour statements fanned out to cluster readers, and rows merged from them.
scramdb_vector_distributed_redispatches_totalcounterReaders of a distributed nearest-neighbour search whose buckets another replica served, after the reader missed its deadline or could not be reached.
scramdb_vector_distributed_failures_totalcounterDistributed nearest-neighbour searches that failed because a reader failed or no replica of a bucket was left.
scramdb_vector_builds_running / scramdb_vector_build_secondsgauge / histogramVector index builds and rebuilds running, and how long one takes.
scramdb_vector_reconcile_rows_totalcounterRows replayed into vector indexes at startup to catch up with the table.
scramdb_vector_build_owed_deletes_totalcounterDeleted rows an index build kept for read snapshots older than the delete, so those snapshots still find them.
scramdb_vector_delete_tail_reclaims_totalcounterReclaims that rebuilt a vector index's delete tail (the deleted rows kept for older snapshots) once some of those rows were no longer needed.
scramdb_vector_delete_tail_batches_dropped_total / scramdb_vector_delete_tail_batches_relinked_totalcounterDelete tail batches a reclaim dropped whole because no read snapshot needs them, and batches it kept whole without copying them.
scramdb_vector_delete_tail_batches_merged_total / scramdb_vector_delete_tail_boundary_batches_folded_totalcounterSmall delete tail batches a reclaim copied into a larger one, and batches it read row by row because a read snapshot's boundary cut through them. Every batch a reclaim judges lands in exactly one of these four counters.
scramdb_vector_fit_sample_bytes_totalcounterBytes of training samples vector index model fits charged to execution memory: a build's fits through storage.memory.maintenance_bytes, every other fit (a streaming INSERT or COPY that trains the model, a background retrain) to the execution pool directly.
scramdb_vector_setting{key} / scramdb_vector_setting_source{key,source}gaugeThe value in force for a [vector] setting, and which layer set it (config, override, setpoint).
scramdb_vector_seriesgaugePrometheus series the vector index families currently emit, computed from the same scrape it appears in.

Per index (labels index, table, column, method, group):

MetricTypeMeaning
scramdb_vector_index_state{state}gaugeWhether the index is in this state (building, ready, regrowing, rebuilding, tail, invalid); one state reads 1.
scramdb_vector_index_live_vectors / scramdb_vector_index_tombstonesgaugeVectors the index holds and can return, and deleted vectors it still carries until compaction.
scramdb_vector_index_unindexed_rowsgaugeRows served exactly because the index has no memory left to hold them (the tail state).
scramdb_vector_index_ledger_live_rows / scramdb_vector_index_delete_tail_rowsgaugeRows the index's row map holds that are not deleted (what a rebuild sizes the index for), and deleted rows the index keeps for read snapshots older than the delete until a reclaim finds no snapshot needs them.
scramdb_vector_index_arena_slots / scramdb_vector_index_arena_used_permillegaugeSlots the index's arena is promised, and the share in use, in thousandths.
scramdb_vector_index_regrows_total / scramdb_vector_index_regrow_runningcounter / gaugeTimes the arena was reopened with a larger promise, and whether it is happening now.
scramdb_vector_index_memory_budget_bytes / scramdb_vector_index_memory_in_use_bytesgaugeThe memory ceiling the index runs under, and resident memory it uses right now.
scramdb_vector_index_memory_plane_bytes{plane}gaugeResident memory by plane: codes, graph, exact_vectors, directory, slot_state, coarse_model, journal_buffers.
scramdb_vector_index_ledger_bytes / scramdb_vector_index_disk_bytes / scramdb_vector_index_journal_bytesgaugeMemory the row-id ledger uses, bytes the index occupies on disk, and bytes of write history kept before compaction.
scramdb_vector_index_ops_total{op} / scramdb_vector_index_ops_failed_total{op}counterOperations completed / refused or failed, op one of search, insert, update, delete, flush.
scramdb_vector_index_search_qpsgaugeSearches per second over the recent window.
scramdb_vector_index_search_latency_seconds{quantile} / scramdb_vector_index_write_latency_seconds{quantile}gaugeSearch / write latency over the recent window, quantile one of 0.5, 0.95, 0.99, 0.999, max.
scramdb_vector_index_stalenessgaugeWrites accepted but not yet searchable. Reads 0: a write is searchable by the very next query, so there is nothing for this gauge to ever report as pending.
scramdb_vector_index_compactions_total / scramdb_vector_index_compaction_runningcounter / gaugeCompactions completed, and whether one is running now.
scramdb_vector_index_fits_totalcounterCoarse model fits the index completed.
scramdb_vector_index_maintenance_running / scramdb_vector_index_graph_armedgaugeWhether the index's own maintenance is running, and whether its navigation graph is built and in use.
scramdb_vector_index_edge_retries_totalcounterNavigation graph edge publications retried because another writer changed the same edge list first. A retry never loses the edge; a high rate means many threads link into the same nodes at once.
scramdb_vector_index_unplaced_slotsgaugeVectors the index's last model fit could not place in the partition it chose for them, for want of room there. While that is a material share of the index, the navigation graph stays off.
scramdb_vector_index_io_reads_total / scramdb_vector_index_io_bytes_read_totalcounterReads and bytes read for vectors spilled to disk.
scramdb_vector_index_cache_hit_ratiogaugeShare of vector reads served from memory.
scramdb_vector_index_simd_level{level}gaugeThe instruction set the index runs with; the matching level reads 1.
scramdb_vector_index_knob{knob}gaugeThe search knob value in force for the index.
scramdb_vector_index_estimator_agreementgaugeShare of candidates whose estimated order matched the exact order in recent searches. Emitted only once a real measurement exists.
scramdb_vector_index_recall{k}gaugeRecall measured by the last recall job for this index. Absent until a recall job has actually run, never a guessed value.
scramdb_vector_index_last_build_wall_microseconds / scramdb_vector_index_last_build_cpu_microsecondsgaugeWall time and CPU time (summed over every thread that worked on it) of the index's last build in this process. Absent until the index is built here: a restart reads absent until the next build, never a stale or guessed figure.
scramdb_vector_index_last_build_peak_memory_bytes / scramdb_vector_index_last_build_rowsgaugeResident memory of the index at the end of its last build (its peak: a build only adds), and rows it indexed.
scramdb_vector_index_search_secondshistogramTime one search of the index took, end to end.
scramdb_vector_index_last_flush_unix_seconds / scramdb_vector_index_rows_since_flushgaugeWhen the index last made its writes durable, and rows written since then.
scramdb_vector_index_last_reindex_unix_secondsgaugeWhen the index was last rebuilt. Emitted only once a REINDEX has actually run.

Cluster (labels index, table, group; present only once a distributed vector index part exists):

MetricTypeMeaning
scramdb_vector_group_owner_is_selfgaugeWhether this node owns the bucket group's index part.
scramdb_vector_group_part_state{state}gaugeThe state of the index part for the bucket group (building, ready, regrowing, rebuilding, tail, invalid, absent, waiting_memory: a hosted bucket group this pass's own memory admission left out of the count it admitted to build).
scramdb_vector_group_rebuilds_total / scramdb_vector_group_rebuild_runningcounter / gaugeTimes the index part was rebuilt after the bucket group moved, and whether one is running now.
scramdb_vector_group_learner_readygaugeWhether a learner may serve vector searches for the bucket group.
scramdb_vector_group_memory_in_use_bytesgaugeResident memory the index part for the bucket group uses.
scramdb_vector_cluster_parts / scramdb_vector_cluster_parts_missinggaugeIndex parts this node holds for distributed tables, and bucket groups it owns whose part is not built.
scramdb_vector_cluster_min_recall{index,k}gaugeThe lowest recall any of this node's parts of the index measured at its last recall run. Absent until measured; the whole cluster's lowest is the lowest over every node.

Recall job (present only on a cluster node, where the job runs):

MetricTypeMeaning
scramdb_vector_recall_parts_measured_total / scramdb_vector_recall_parts_skipped_totalcounterIndex parts the recall job measured on this node, and parts a pass was due to measure and could not (each such part says why in scram_vector_index_parts.recall_note).
scramdb_vector_recall_pass_cpu_secondshistogramCPU time of one recall pass on this node.
scramdb_vector_recall_next_wait_seconds{reason}gaugeSeconds until the job next measures a part, and why: churn (the wait follows how much of a part changed), cost (the one-percent CPU duty cycle binds), ceiling (a lightly changed part waits the hour ceiling), pinned (vector.recall_interval_secs), first (a part never measured), unchanged (nothing changed since the last measurement, so nothing is due: +Inf), or off.
scramdb_vector_recall_probes{reason}gaugeProbe queries the last measurement sampled, and why: precision (enough for a one-point standard error at the part's last recall), floor (8), ceiling (64), memory (the vector memory budget held fewer), or pinned (vector.recall_probes).

Vector index system views​

The same sample backs five read-only system views, so a dashboard that prefers SQL over scraping Prometheus sees the identical numbers. Every row leads with node, the node that measured it.

On a cluster, the four statistics views answer for the whole cluster from whichever node you query: that node's rows and every other member's, gathered at query time. A member that does not answer within vector.stats_node_timeout_ms (default 5000) contributes one row with its node, the word unreachable in state (or metric), the reason in error, and NULL everywhere else, never a silently shorter result. scram_vector_settings is the queried node's own settings. On a single node every view is that node's rows.

ViewColumns
scram_vector_settingsnode, key, value, source, unit, live, description - every [vector] key, the value in force, and which layer set it
scram_vector_indexesnode, index, schema, table, column, method, metric, opclass, state, arena_slots, arena_used_permille, regrows, live, tombstones, unindexed_rows, memory_budget, memory_in_use, disk_bytes, journal_bytes, search_qps, p50_us, p99_us, staleness, compactions, fits, graph_armed, simd_level, estimator_agreement, recall_at_10, last_flush, last_reindex, options, build_wall_us, build_cpu_us, build_peak_memory, build_rows, search_count, search_p99_us, error - one row per index per node; a distributed index's parts on a node fold into its row (build cost summed, NULL unless every part was built by that process; recall the lowest part's)
scram_vector_index_partsnode, index, schema, table, group, owner, is_self, state, live, memory_in_use, rebuilds, learner_ready, rebuild_running, recall_at_10, recall_probes, recall_measured_at, recall_note, error - one row per bucket group each node hosts; is_self marks the rows of the node you queried; recall_note says why a part has no recall figure (too few vectors, over vector.recall_max_rows, not ready, or no memory); empty on a single node
scram_vector_memorynode, total, reserved, in_use, engine_bytes, refusals, error - one summary row per node
scram_vector_activitynode, metric, value, unit, description, error - the node-wide vector counters, one row per counter per node
SELECT node, index, state, live, memory_in_use, search_qps FROM scram_vector_indexes;
SELECT node, "group", state, live, recall_at_10, recall_note
FROM scram_vector_index_parts WHERE index = 'docs_embedding_idx' OR error IS NOT NULL;
SELECT key, value, source, live FROM scram_vector_settings WHERE key = 'vector.over_fetch';

EXPLAIN on an indexed vector query names the tier chosen and the candidate count; EXPLAIN ANALYZE adds candidates fetched, widening rounds taken, re-rank time, and whether the answer came back underfilled. See Vector Search for the query shapes these views and counters describe, and Tuning: Vector Indexes for how to act on them.

What is not on this endpoint​

  • GPU metrics are not exposed today. There is no scramdb_gpu_* family on /metrics, even though [gpu] metrics_enabled defaults to true. See GPU Acceleration for the log-based way to confirm GPU dispatch instead.
  • No buffer-pool hit-ratio, cache-hit-ratio, or per-worker-utilization metric exists. Memory and parallelism verification on this site relies on \timing and OS-level observation instead, see Memory and Parallelism.

Log levels and format​

SCRAMDB_LOG=debug SCRAMDB_LOG_FORMAT=json scramdb -c scramdb-config.toml --pg-address 127.0.0.1:5432 --pg-no-auth

Expected: stderr lines shaped like {"ts":"...","level":"INFO","target":"...","msg":"Metrics endpoint listening on http://0.0.0.0:9090/metrics"}.

SCRAMDB_LOG​

Controls the level filter. Accepted values (case-insensitive): off, error, warn, info, debug, trace. Default when unset: warn, so a release binary starts quiet. An unrecognized value does not fail loud, it silently falls back to warn, so a typo in this env var will not show up as an error, it just quietly serves you the default level.

SCRAMDB_LOG_FORMAT​

Set to json (case-insensitive) for structured JSON output. Anything else, including leaving it unset, produces the default plain-text format.

Plain-text format: [{level}] [{timestamp}] {message}, for example:

[INFO] [2026-08-02T10:15:30.123Z] Metrics endpoint listening on http://0.0.0.0:9090/metrics

JSON format: {"ts":"<timestamp>","level":"<LEVEL>","target":"<module path>","msg":"<message>"}. Only backslash (\) and double-quote (") characters in the message are escaped. No other JSON escaping is performed, so a log message containing a raw newline or control character is not escaped and can produce an invalid JSON line in that rare case. Keep this in mind if you feed the JSON log output into a strict JSON-line parser downstream.

Logging is asynchronous: a dedicated low-priority background thread drains log messages and writes to stderr, so logging never blocks a query worker. Output always goes to stderr, never directly to a file, redirect it yourself (scramdb ... 2> scramdb.log) or capture it with your process supervisor.

One additional env var is useful for deeper diagnosis alongside SCRAMDB_LOG, covered on another page in this section:

  • SCRAMDB_GPU_FORCE=1, forces GPU pipeline-shape eligibility for benchmarking. See GPU Acceleration.

Next​

See Query Performance for how the JIT chunk counters are used in practice, GPU Acceleration for why GPU verification uses logs instead of this endpoint, or Clustering for how the cluster counters above map onto a Docker Compose or Kubernetes deployment.