Observability
By the end of this page you will know how to scrape ScramDB's metrics, check liveness, and configure logging, with the exact set of metric names that exist today so you do not build a dashboard around a counter that isn't there.
The metrics and health endpoint
ScramDB serves both /metrics and /health on the same HTTP listener, one port, one process.
| Setting | Default | How to set it |
|---|---|---|
| Port | 9090 | [general] metrics_port in TOML, or --metrics-port <PORT> on the command line. Set to 0 to disable the endpoint entirely. |
| Bind address | 0.0.0.0 (all interfaces) | Not configurable. If you run ScramDB without a firewall or network policy in front of it, this port is reachable from outside the host. |
/metrics
curl -s http://127.0.0.1:9090/metrics | head -20
Expected output: Prometheus text exposition format (Content-Type: text/plain; version=0.0.4; charset=utf-8), starting with lines like # HELP scramdb_queries_total .... If it fails: connection refused means either the server isn't running, metrics_port is set to 0, or a firewall is blocking the port. Remember the listener binds 0.0.0.0, so "connection refused from outside the host" is more likely a firewall rule than the endpoint being loopback-only.
Any path other than /metrics or /health returns 404 Not Found.
/health
curl -s http://127.0.0.1:9090/health
Expected output: ok (plain text, no JSON, no fields, HTTP 200). On a cluster node whose commit protocol has stopped after a failure (scramdb_cluster_dilith_host_stopped at 1), the answer is HTTP 503 Service Unavailable with the body the commit protocol stopped on this node: that node can no longer commit transactions and must be restarted, so a load balancer or orchestrator probing /health takes it out of service. There is no separate liveness/readiness distinction. Beyond that one case it does not check database or storage health internally: a 200 ok proves the process and its HTTP listener are up, not that queries are succeeding. For query success/failure, use the scramdb_queries_ok / scramdb_queries_err counters on /metrics instead.
The endpoint reports on itself on the same /metrics body, always emitted:
| Metric | Type | Meaning |
|---|---|---|
scramdb_metrics_endpoint_requests_total | counter | Requests the metrics endpoint answered. |
scramdb_metrics_endpoint_request_timeouts_total | counter | Clients disconnected because their request did not arrive in time. |
scramdb_metrics_endpoint_response_timeouts_total | counter | Clients disconnected because they did not take their response in time. |
scramdb_metrics_endpoint_client_errors_total | counter | Client connections that failed before their answer was written. |
scramdb_metrics_endpoint_accept_failures_total | counter | Connections the metrics endpoint could not accept. |
scramdb_metrics_endpoint_render_seconds | histogram | Time to render one /metrics answer. |
Metric reference
Every metric named on this page is really emitted on /metrics, so you can build a dashboard from these names without checking first. The reverse does not hold: the endpoint also carries deep diagnostic counters that are not part of the supported surface and can change between releases. The families below are the ones worth alerting on and graphing.
Core query and connection metrics (always present)
| Metric | Type | Meaning |
|---|---|---|
scramdb_queries_total | counter | Total queries received. |
scramdb_queries_ok | counter | Queries that completed successfully. |
scramdb_queries_err | counter | Queries that errored. |
scramdb_queries_active | gauge | Queries currently executing. |
scramdb_query_avg_us | gauge | Rolling average query latency in microseconds. |
scramdb_connections_total | counter | Total connections accepted since startup. |
scramdb_connections_active | gauge | Currently open connections. |
scramdb_bytes_sent | counter | Total bytes sent to clients. |
scramdb_bytes_received | counter | Total bytes received from clients. |
scramdb_uptime_seconds | gauge | Seconds since the process started. |
Execution-tier metrics (always present)
| Metric | Type | Meaning |
|---|---|---|
scramdb_chunks_jit_total | counter | Morsel-chunks executed as compiled JIT machine code. |
scramdb_chunks_vm_total | counter | Morsel-chunks executed by the bytecode interpreter. |
scramdb_compiled_calls_total | counter | Sub-region JIT invocations. |
These are the counters used in the JIT scrape-delta technique described on Query Performance.
Client connections (always present)
Every client connection holds one of [general] max_connections places while its session lives, the last superuser_reserved_connections of them kept for superusers. An idle session waits for its client without holding a thread (it is parked) and is handed back to a thread when its client sends something or a deadline passes.
| Metric | Type | Meaning |
|---|---|---|
scramdb_max_connections | gauge | Connection places the client listener enforces (max_connections). |
scramdb_connection_places_taken | gauge | Connection places held by sessions. |
scramdb_connection_starting | gauge | Connections still negotiating SSL, sending their startup packet or authenticating. |
scramdb_connection_parked | gauge | Idle sessions waiting for their client with no thread held. |
scramdb_connection_worker_threads | gauge | Threads that serve client sessions, busy or cached. |
scramdb_connection_idle_workers | gauge | Session threads cached with no session to run. |
scramdb_connection_wakes_total | counter | Parked sessions handed to a thread because their client sent data or a deadline passed. |
scramdb_connection_refused_too_many_total | counter | Connections refused because every connection place was taken. |
scramdb_connection_refused_reserved_total | counter | Connections refused because only the places reserved for superusers were left. |
scramdb_connection_refused_role_limit_total | counter | Connections refused by their role's connection limit. |
scramdb_connection_refused_server_setting_total | counter | Connections refused because they asked to change a setting fixed when the server started (max_connections or superuser_reserved_connections in their startup options, FATAL 55P02). |
scramdb_connection_refused_startup_setting_total | counter | Connections refused because a setting their startup packet sent (TimeZone, DateStyle or any other) had a value the setting refuses (FATAL, the setting's own code, as PostgreSQL). |
scramdb_connection_hba_rejects_total | counter | Connections rejected by the host-based authentication rules when accepted. |
scramdb_connection_idle_session_timeouts_total | counter | Sessions closed because they stayed idle outside a transaction past idle_session_timeout (FATAL 57P05). |
scramdb_connection_startup_timeouts_total | counter | Connections closed because authentication did not finish within authentication_timeout. |
scramdb_connection_tls_failures_total | counter | Client TLS handshakes that failed. |
scramdb_connection_messages_refused_total | counter | Client messages refused because the server could not reserve memory to receive them. |
scramdb_connection_worker_start_failures_total | counter | Session threads that could not be started. |
scramdb_connection_ended_client_gone_total | counter | Sessions whose client closed its connection without a Terminate message. |
scramdb_connection_ended_socket_error_total | counter | Sessions whose client connection failed. |
scramdb_connection_ended_terminated_total | counter | Sessions closed by pg_terminate_backend or a server shutdown. |
scramdb_connection_ended_in_transaction_total | counter | Sessions that ended with a transaction open, which was rolled back (its row, table and advisory locks released with it). |
scramdb_connection_startup_seconds | histogram | Time from accepting a connection to its session being ready for queries. |
refused_too_many_total climbing means clients open more connections than max_connections allows: put a pooler in front, or raise the limit. idle_session_timeouts_total moves only when a session sets idle_session_timeout (off by default).
Statements (always present)
A session hands each statement to a query worker and waits for its result; a statement timeout, a cancel, the transaction timeout and a lost client all end that wait at once.
| Metric | Type | Meaning |
|---|---|---|
scramdb_statement_wait_seconds | histogram | Time a session waited for its statement's result from a query worker. |
scramdb_statement_timeouts_total | counter | Statements cancelled because they ran past statement_timeout. |
scramdb_statement_cancels_total | counter | Statements a client, a forced close or the memory watchdog cancelled while they ran. |
scramdb_statement_transaction_timeouts_total | counter | Sessions closed because their transaction ran past transaction_timeout during a statement. |
scramdb_statement_clients_gone_total | counter | Sessions closed because their client went away during a statement. |
scramdb_statement_late_results_dropped_total | counter | Results of statements a session had already stopped waiting for, dropped when they arrived. |
scramdb_statement_implicit_blocks_total | counter | Multi-statement queries and pipelines run as one implicit transaction, as PostgreSQL runs them. |
scramdb_statement_implicit_blocks_rolled_back_total | counter | Implicit transactions rolled back because one of their statements failed. |
Prepared statements (always present)
A prepared statement is parsed once, the first time a Bind needs it; every later execution hands the engine that parse with the Bind's values in place, so a statement executed many times is never parsed again.
| Metric | Type | Meaning |
|---|---|---|
scramdb_prepared_template_parses_total | counter | Prepared statements parsed, once each, the first time a Bind needed the parse. |
scramdb_prepared_handoffs_total | counter | Prepared statement executions handed to the engine already parsed. |
scramdb_prepared_handoffs_unused_total | counter | Prepared statement executions whose parse the engine did not use. |
scramdb_prepared_reparses_total | counter | Times the engine parsed an already parsed prepared statement execution again. |
On a workload of prepared statements, handoffs_total grows with the executions while template_parses_total grows only with the distinct statements; handoffs_unused_total and reparses_total staying near zero confirm the parse is reused.
Plan cache (always present)
Plans, their compiled code and plan markers are cached per statement and bounded in bytes by [optimizer] plan_cache_memory.
| Metric | Type | Meaning |
|---|---|---|
scramdb_plan_cache_hits_total | counter | Statements that reused a cached plan. |
scramdb_plan_cache_misses_total | counter | Statements that found no current cached plan and were planned. |
scramdb_plan_cache_evictions_total | counter | Cached plans and plan markers dropped to stay within the plan cache memory budget. |
scramdb_plan_cache_bytes | gauge | Memory held by cached plans, their compiled modules and plan markers. |
scramdb_plan_cache_budget_bytes | gauge | Memory the plan cache may hold. |
scramdb_plan_cache_entries | gauge | Plans and plan markers the plan cache holds. |
Notifications (LISTEN and NOTIFY)
LISTEN and NOTIFY report process-wide, always emitted:
| Metric | Type | Meaning |
|---|---|---|
scramdb_notify_notifications_committed_total | counter | Notifications sent by committed transactions. |
scramdb_notify_notifications_delivered_total | counter | Notifications written to listening clients. |
scramdb_notify_batches_replicated_total | counter | Committed transactions whose notifications were sent to the other nodes. |
scramdb_notify_replication_failures_total | counter | Committed transactions whose notifications could not be sent to the other nodes. |
scramdb_notify_queue_full_total | counter | Notifications refused because the notification queue was full. |
scramdb_notify_listening_sessions | gauge | Sessions listening to at least one channel. |
scramdb_notify_held_bytes | gauge | Bytes of notifications queued or waiting to be sent to a listener. |
A notification's bytes are held from the execution memory budget until the last listener has received it, so held_bytes that keeps growing means a session listens and never lets its client read (a client blocked inside a long transaction receives nothing until it ends); queue_full_total counts the NOTIFY statements refused with 54000 when that budget had no room. On a cluster, replication_failures_total counts committed transactions whose notifications did not reach the other nodes: the committing client got a warning (08006) with its commit, and listeners on other nodes missed those notifications.
WAL archive metrics
| Metric | Type | Meaning |
|---|---|---|
scramdb_wal_archive_enabled | gauge | 1 if WAL archival is enabled, 0 otherwise. Always emitted. |
scramdb_wal_archive_errors_total | counter | Archive attempts that failed. Always emitted. |
scramdb_wal_archive_segments_archived_total | counter | Segments successfully shipped to the destination. Always emitted. |
scramdb_wal_archive_segments_behind | gauge | Segments not yet archived. Emitted only once a real measurement exists. |
scramdb_wal_archive_bytes_behind | gauge | Bytes not yet archived. Emitted only once a real measurement exists. |
scramdb_wal_archive_last_archived_lsn | gauge | LSN of the last successfully archived segment. Emitted only once a real measurement exists. |
scramdb_wal_archive_segments_pruned_total | counter | WAL segments removed past the archive retention window ([storage.wal.archive] retention). Always emitted. |
scramdb_wal_archive_prune_failures_total | counter | Archive retention passes that failed and will retry. Always emitted; zero on a healthy archive. |
The periodic checkpoint scheduler ([storage.wal.checkpoint], on by default) reports on the same endpoint, always emitted:
| Metric | Type | Meaning |
|---|---|---|
scramdb_wal_checkpoint_interval_total | counter | Checkpoints the scheduler ran because its interval elapsed. |
scramdb_wal_checkpoint_wal_bytes_total | counter | Checkpoints the scheduler ran because the WAL grew past its threshold. |
scramdb_wal_checkpoint_behind_total | counter | Checkpoints run at full disk priority because the WAL grew past twice its threshold. |
scramdb_wal_checkpoint_failures_total | counter | Checkpoints the scheduler started that failed. |
scramdb_wal_checkpoint_threshold_bytes | gauge | WAL bytes written since the last checkpoint that start the next one. |
scramdb_wal_checkpoint_seconds | histogram | Time one scheduled checkpoint took. |
The cold tier (Storage Tiering) reports on the same endpoint, always emitted; every failure counter stays at 0 on a healthy tier:
| Metric | Type | Meaning |
|---|---|---|
scramdb_cold_segments_demoted_total | counter | Sealed segments copied to cold storage. |
scramdb_cold_bytes_demoted_total | counter | Bytes of sealed segments copied to cold storage. |
scramdb_cold_segments_evicted_total | counter | Segments whose local file was removed because cold storage holds them. |
scramdb_cold_bytes_evicted_total | counter | Bytes of local disk freed by segments cold storage holds. |
scramdb_cold_segments_read_back_total | counter | Segments read back from cold storage to local disk. |
scramdb_cold_bytes_read_back_total | counter | Bytes read back from cold storage to local disk. |
scramdb_cold_demotion_failures_total | counter | Segment copies to cold storage that failed and will be tried again. |
scramdb_cold_read_back_failures_total | counter | Segment reads from cold storage that failed after every attempt. |
scramdb_cold_read_back_retries_total | counter | Segment reads from cold storage attempted again after a failure. |
scramdb_cold_read_back_mismatches_total | counter | Cold storage objects that did not hold the bytes their segment's marker names. |
scramdb_cold_transfer_timeouts_total | counter | Transfers to or from cold storage that ran past their deadline. |
Operators past their memory (work_mem, see Memory) report their spills, always emitted:
| Metric | Type | Meaning |
|---|---|---|
scramdb_sort_spill_runs_total | counter | Sorted runs written to disk by sorts past their memory. |
scramdb_sort_spill_bytes_written_total | counter | Bytes of sorted runs written to disk. |
scramdb_join_build_spill_partitions_total | counter | Hash join build partitions written to disk to stay within their memory. |
The compiled-function cache of [udf] (aot_cache_dir, aot_cache_max_bytes) reports on the same endpoint:
| Metric | Type | Meaning |
|---|---|---|
scramdb_udf_compile_cache_bytes | gauge | Bytes of compiled function modules held in the on-disk compile cache. |
scramdb_udf_compile_cache_limit_bytes | gauge | The most bytes the compile cache holds (aot_cache_max_bytes). |
scramdb_udf_compile_cache_hits_total | counter | Function modules loaded from the compile cache. |
scramdb_udf_compile_cache_misses_total | counter | Function modules compiled because the cache did not hold them. |
scramdb_udf_compile_cache_evictions_total | counter | Modules removed to keep the cache within its bound. |
scramdb_udf_compile_cache_write_failures_total | counter | Compiled modules that could not be saved to the cache. |
scramdb_udf_catalog_saves_total | counter | Saves of the installed modules and functions, one per INSTALL MODULE, DROP MODULE, CREATE FUNCTION and DROP FUNCTION of a module-backed language. |
scramdb_udf_catalog_save_failures_total | counter | Module and function changes undone because they could not be saved (the statement failed with 58030). |
scramdb_udf_catalog_modules_loaded_total | counter | Installed modules loaded at startup. |
scramdb_udf_catalog_functions_loaded_total | counter | Functions loaded at startup. |
scramdb_udf_catalog_load_failures_total | counter | Saved modules and functions that could not be loaded at startup; the server log names each. |
scramdb_udf_catalog_module_files_removed_total | counter | Saved module files removed because no module uses them any more. |
scramdb_udf_epoch_ticks_total | counter | Ticks of the function engine's deadline clock, which ticks only while a function runs. |
scramdb_udf_epoch_ticker_wakes_total | counter | Times the function engine's deadline clock woke for a call after being idle. |
The WAL ring reports on the same endpoint, always emitted. The bound is what the memory budget reserves for it: the ring, 64 times [storage.wal] flush_threshold_bytes (sized for the machine unless pinned), and its completion slots (a sixteenth of the ring). append_waits_total rising means the WAL device is not keeping up with the write rate. A stall is the log making no progress for far longer than a flush usually takes (at least a second at the default epoch): it is reported, never treated as a failure, and every commit waiting for it completes when the device answers. An interrupted log write leaves padding in its place, which every reader skips:
| Metric | Type | Meaning |
|---|---|---|
scramdb_wal_buffer_bytes | gauge | Bytes of log records waiting in the WAL ring at its last flush. |
scramdb_wal_buffer_bound_bytes | gauge | Bytes the WAL ring and its completion slots hold, as the memory budget reserves them. |
scramdb_wal_append_waits_total | counter | WAL appends that waited for a flush to make room in the ring. |
scramdb_wal_flush_threshold_bytes | gauge | Bytes of waiting log records that start a WAL flush early, as configured or sized. |
scramdb_wal_flush_stalls_total | counter | Times the WAL made no progress for far longer than a flush usually takes. |
scramdb_wal_flush_stalled | gauge | WAL writers whose log has made no progress for far longer than usual right now. |
scramdb_wal_padding_records_total | counter | Log records whose append was interrupted and whose place was filled with padding. |
scramdb_wal_flush_syncs_total | counter | Data syncs of the WAL that flush rounds made, one per round that wrote log records. |
scramdb_wal_rotation_syncs_total | counter | Data and directory syncs of the WAL made when a file filled and the next one began. |
scramdb_wal_flush_sync_seconds | histogram | Time one flush round's WAL data sync took. |
Every sync of the WAL is counted: the sum of flush_syncs_total and rotation_syncs_total over the transactions committed in the same window is the syncs per committed transaction, and flush_sync_seconds is how long the device takes to make one flush round durable.
Catalog file writes (each table's catalog file and deleted-row record, the table list, sequences, databases and the commit timestamp mark) report on the same endpoint, always emitted. Concurrent changes to one file share one write, so coalesced_total rising while writes_total stays flat under heavy DDL or checkpointing is expected:
| Metric | Type | Meaning |
|---|---|---|
scramdb_catalog_writes_total | counter | Catalog files written. |
scramdb_catalog_write_failures_total | counter | Catalog file writes that failed. Zero on a healthy disk; each failure is returned to the statements that asked for the write. |
scramdb_catalog_coalesced_total | counter | Catalog changes made durable by a write another caller did. |
scramdb_catalog_write_seconds | histogram | Time one catalog file write took (write, sync, rename, directory sync). |
The segments_behind, bytes_behind and last_archived_lsn metrics being absent from a scrape means "not measured yet," not zero. This is deliberate: the endpoint does not fabricate a 0 for a value it has not actually computed. Treat an absent metric name the same way, as unmeasured, everywhere on this endpoint, not just for WAL archive.
Object store and function IO threads
The object store client (the s3://, gs:// and az:// destinations of the WAL archive, base backups, restores and branch staging) and the file, clock and network calls user functions make each run on a small pool of worker threads of their own. A pool starts with the first request that needs it, so a server whose destinations are local directories, or that runs no function, never starts one. Always emitted: zero until the pool starts.
| Metric | Type | Meaning |
|---|---|---|
scramdb_object_store_io_workers | gauge | Worker threads of the object store client. |
scramdb_object_store_io_busy_microseconds_total | counter | Time the object store client's worker threads spent running work. |
scramdb_udf_io_workers | gauge | Worker threads serving functions' file, clock and network calls. |
scramdb_udf_io_busy_microseconds_total | counter | Time the worker threads serving functions' IO calls spent running work. |
A busy counter's rate divided by its worker gauge is that pool's utilization.
Every upload and download through the object store client is counted, and so is every HTTP read of a remote file. A request is bounded: 5 s to connect and 30 s of silence. A read that got no answer (a connection that did not open or broke, a name that did not resolve, a timeout, or a 408, 429, 500, 502, 503 or 504) is sent again, up to 10 times within 180 s, backing off from 100 ms to 15 s; an object store request keeps the same bounds. Always emitted: zero until the first request.
| Metric | Type | Meaning |
|---|---|---|
scramdb_object_store_uploads_total / scramdb_object_store_upload_failures_total | counter | Uploads to the object store, and the ones that failed. |
scramdb_object_store_uploads_aborted_total / scramdb_object_store_abort_failures_total | counter | Unfinished multipart uploads aborted after a failure, and the aborts that failed (the object store keeps those parts until its own cleanup). |
scramdb_object_store_uploaded_bytes_total | counter | Bytes uploaded. |
scramdb_object_store_downloads_total / scramdb_object_store_download_failures_total | counter | Downloads from the object store, and the ones that failed. |
scramdb_object_store_downloaded_bytes_total | counter | Bytes downloaded. |
scramdb_object_store_upload_seconds / scramdb_object_store_download_seconds | histogram | Time an upload or a download took. |
scramdb_remote_file_requests_total / scramdb_remote_file_request_failures_total | counter | HTTP requests for remote files, and the ones that failed after their retries. |
scramdb_remote_file_request_retries_total | counter | HTTP requests for remote files sent again. |
scramdb_remote_file_request_timeouts_total | counter | HTTP requests for remote files that got no answer within their bounds. |
Resilience metrics (always present, all zero on a healthy process)
| Metric | Type | Meaning |
|---|---|---|
scramdb_parallel_join_timeout_total | counter | Parallel join operations that hit their timeout. |
scramdb_pool_watchdog_heals_total | counter | Buffer pool watchdog self-heal actions taken. |
scramdb_pool_watchdog_heal_budget_exhausted_total | counter | Times the watchdog's heal budget ran out. |
scramdb_pool_watchdog_escalations_total | counter | Times the watchdog escalated beyond a self-heal. |
Query execution and JIT
| Metric | Type | Meaning |
|---|---|---|
scramdb_jit_compile_submits_total | counter | Compilation requests handed to the JIT worker. On a warm, repeating workload this should stop growing; continued growth means the compiled-artifact cache is not serving and is worth raising compiled_cache_memory for. |
scramdb_chunks_jit_t1_total | counter | Chunks served by tier 1 compiled code, which is the fast-to-produce tier or a cached artifact loaded from disk. |
scramdb_chunks_jit_t2_total | counter | Chunks served by tier 2 compiled code, the fully optimized tier. |
scramdb_jit_t1_compile_seconds | histogram | Time to compile one module for a quick start (tier 1). |
scramdb_jit_t2_compile_seconds | histogram | Time to compile one module fully optimized (tier 2). |
scramdb_jit_jetz_load_seconds | histogram | Time to load one compiled module from the on-disk compile cache. |
scramdb_jit_jetbc_load_seconds | histogram | Time to load and compile one cached intermediate module from the on-disk compile cache. |
scramdb_jit_verifier_rejections_total{tier} | counter | Compiled kernels the code verifier rejected, labelled tier (fused, whole_function, expression_region, module). The statement still runs correctly on the interpreter; a nonzero count is a code generator defect worth reporting, not a query problem. |
scramdb_udf_interpreter_warmups_total | counter | Interpreted function runtimes started, by language (python, ruby): each starts once per process, the first time a function of that language runs. |
scramdb_jit_compile_memory_charges_total | counter | Compiles whose working memory was charged to the compiled-code cache budget. |
scramdb_jit_compile_memory_in_flight_bytes | gauge | Working memory held by the compiles running now. It counts against compiled_cache_memory together with the cached code, and cached code is evicted to make room for it. |
scramdb_jit_compiles_in_flight | gauge | Compiles running now. |
scramdb_jit_compile_memory_over_budget_total | counter | Compiles that started while the running compiles alone needed more than compiled_cache_memory. A compile is never refused or delayed for memory, so a growing count means the budget is too small for the compile concurrency: raise compiled_cache_memory or lower jit_compile_threads. |
scramdb_jit_compile_threads | gauge | Threads of the query compile pool ([execution] jit_compile_threads, one to eight). They run on the query cores below the query threads' priority, never on the cores a cluster node reserves for its own traffic. |
scramdb_jit_compile_queued | gauge | Query compiles waiting for a compile thread. |
scramdb_jit_compile_queued_bytes | gauge | Bytes held by the query compiles waiting for a compile thread. Bounded by compiled_cache_memory. |
scramdb_jit_compile_refused_total | counter | Query compiles not queued because the compile queue was full. The statement runs on the interpreter, never waits, and is compiled on a later run once the queue has room; a steadily growing count means compiles cannot keep up with new statement shapes: raise jit_compile_threads or compiled_cache_memory. |
scramdb_vm_agg_flushes_total | counter | Interpreter-side grouped aggregate tables flushed as partial results on reaching their memory budget. |
scramdb_preagg_flushes_total | counter | Per-worker aggregate tables flushed as partial results for exceeding their share of the execution budget. |
scramdb_preagg_merges_on_caller_total | counter | Small pre-aggregate merges run on the calling thread. |
scramdb_preagg_merge_empty_lanes_total | counter | Empty pre-aggregate lanes skipped at merge. |
scramdb_agg_resident_bytes | gauge | Live aggregate hash-table bytes: buckets, key slots, tags and string arenas. |
scramdb_join_ordering_cartesian_fallbacks_total | counter | Times the planner could not connect a join graph and fell back to emitting a cross product. Any query whose SQL does not ask for a cross join should leave this at zero; a non-zero delta is worth reporting. |
scramdb_split_distinct_rewrites_total | counter | DISTINCT aggregates the planner split into two aggregation levels. |
scramdb_probe_side_materializations_total | counter | Outer-join probe sides that could not stream and were written to a temporary relation first. |
scramdb_join_input_materializations_total | counter | Join inputs of any position that could not stream and were written to a temporary relation first. Falling over time on a fixed workload is the plan getting better. |
scramdb_stream_topk_rows_total | counter | Rows accepted by the streaming top-K sink serving ORDER BY ... LIMIT. |
scramdb_topk_threshold_publishes_total | counter | Top-K thresholds published into the scan so it can skip pages that cannot beat the current best rows. |
scramdb_late_fetch_rows_total | counter | Rows served by late materialization: a narrow scan that fetches the remaining columns only for rows that survived. |
scramdb_stream_queue_bytes | gauge | Bytes of worker results queued ahead of the client, bounded by backpressure. |
scramdb_stream_gate_forced_admits_total | counter | Results pushed past their statement's budget because the queue was empty and there was nothing left to drain. |
Startup
| Metric | Type | Meaning |
|---|---|---|
scramdb_startup_subsystems_pending | gauge | Subsystems a client can reach (vector indexes, the memory watchdog, the function engine) that are not ready yet. The PostgreSQL port accepts no client until this is zero. |
scramdb_startup_listener_wait_micros | gauge | Microseconds the client listener waited for startup to finish before accepting clients. Zero on an ordinary boot, where every subsystem is ready before the port is opened. |
No statement fails because a subsystem is still being installed: a client whose connection reaches the listener before startup has finished waits and is served once it has.
Scans and storage
| Metric | Type | Meaning |
|---|---|---|
scramdb_index_delete_lock_acquisitions_total | counter | Batches of entries removed from ordered and hashed indexes. Index reads never wait on these removals. |
scramdb_index_delete_entries_total | counter | Entries asked to be removed from ordered and hashed indexes. |
scramdb_index_read_gate_raised_total | counter | Index key reads that held the index's writers back to finish: a read of a key that writers kept changing through a few tries pauses every change to that index until it has read the key once. Rare outside a storm of writes on one key. |
scramdb_index_read_gate_writer_waits_total | counter | Index changes that waited for such a key read to finish. Each wait lasts one read of the key. |
scramdb_zone_pages_considered_total | counter | Pages evaluated against zone-map predicates. |
scramdb_zone_pages_pruned_total | counter | Pages skipped by zone-map pruning and never read. Against considered, this is your page-pruning hit rate. |
scramdb_selective_pages_read_total | counter | Pages served by reading only the projected columns rather than the whole page. |
scramdb_selective_bytes_saved_total | counter | Bytes that column-selective reading did not have to fetch. |
scramdb_selective_ranges_read_total | counter | Read requests column-selective reading issued. Divided by selective_pages_read_total it is the number of disk reads per page. |
scramdb_read_pool_active_units | gauge | Units currently holding a read-buffer-pool slot, either waiting for concurrency or actively holding pages. |
scramdb_read_pool_unit_waits_total | counter | Times a unit waited noticeably long for a read-pool concurrency slot. This is a throughput signal, not a safety one: page acquisition itself is never queued. |
scramdb_row_fetches_total | counter | Fetches of rows by their address, for index probes, row re-reads and re-ranking. |
scramdb_row_fetch_rows_total | counter | Rows returned by fetches of rows by their address. |
scramdb_row_fetch_pages_total | counter | Pages read by fetches of rows by their address. |
scramdb_row_fetch_page_bytes_total | counter | Bytes of the pages read by fetches of rows by their address. |
scramdb_row_fetch_decoded_bytes_total | counter | Bytes decoded from those pages by fetches of rows by their address. Against page_bytes_total, how much of each page a fetch had to decode. |
scramdb_values_out_of_line_written_total | counter | Values too long for their page written out of line, in a chain of pages of their own. |
scramdb_value_chain_pages_written_total | counter | Pages written to hold values stored out of line. |
scramdb_values_out_of_line_bytes_written_total | counter | Bytes of the values written out of line. |
scramdb_value_chain_bytes_stored_total | counter | Bytes stored for values written out of line, after compression. Against bytes_written_total, the compression those values got. |
scramdb_values_out_of_line_read_total | counter | Values stored out of line read back. |
scramdb_value_chain_pages_read_total | counter | Pages read to read values stored out of line. |
scramdb_values_out_of_line_bytes_read_total | counter | Bytes of the values stored out of line read back. |
Rows fetched by address are counted by each query worker on its own, so the counters cost nothing measurable however many workers fetch. A text, binary, vector or array value longer than fits its page is stored out of line, in a chain of pages of its row's segment, and read back only by the rows that need it; the out-of-line counters are zero on a workload with no such values.
Snapshot index reads
An index lists only the latest version of each row. A REPEATABLE READ or SERIALIZABLE read whose snapshot is older than a change finds the version it sees through the index's record of superseded entries, instead of scanning the table.
| Metric | Type | Meaning |
|---|---|---|
scramdb_snapshot_index_probes_versioned_total | counter | Snapshot index probes answered from the index and its record of superseded entries. |
scramdb_snapshot_index_probes_scanned_total | counter | Snapshot index probes answered by scanning the table. |
scramdb_retired_versions_recorded_total | counter | Superseded index entries kept for snapshots that can still see them. |
scramdb_retired_versions_released_total | counter | Superseded index entries released by compaction or with their index. |
scramdb_retired_versions_entries | gauge | Superseded index entries currently kept for snapshots. |
probes_scanned_total climbing against probes_versioned_total means snapshot reads are scanning whole tables instead of probing their index.
Read snapshot pins
The snapshot of a statement, of a REPEATABLE READ or SERIALIZABLE transaction, of a streaming result and of a distributed fragment's local scan is pinned while it is in use, so compaction keeps every row version it can still see.
| Metric | Type | Meaning |
|---|---|---|
scramdb_read_pins_taken_total | counter | Read snapshot pins registered. |
scramdb_read_pins_released_total | counter | Read snapshot pins released. |
scramdb_read_pin_redraws_total | counter | Snapshot draws that moved their pin off the floor registered before the draw. |
scramdb_read_pins_held | gauge | Read snapshot pins held now. |
scramdb_retention_holds_taken_total | counter | Retention holds taken for the shard groups hosted on this node. |
scramdb_retention_holds_released_total | counter | Retention holds released. |
scramdb_retention_holds_held | gauge | Retention holds held now for the shard groups hosted on this node. |
scramdb_retention_boot_holds_armed_total | counter | Stores that kept every row version at start until their shard groups were pinned. |
scramdb_retention_boot_holds_released_total | counter | Start-up retention holds released. |
scramdb_retention_boot_holds_held | gauge | Stores keeping every row version until their shard groups are pinned. |
read_pins_held that keeps growing on a steady workload means snapshots are not being released, which holds old row versions on disk; look for long-open transactions or unread result streams.
On a cluster node, each shard group hosted here also holds the row versions between the oldest read position the group can still serve and its newest commit, since a read at such a position can reach this node long after it was chosen; a quiet group holds only its last window, never what busier groups write after it. retention_holds_held follows the groups hosted here. At start a store keeps every row version until each shard group it holds rows of is hosted again or placed on another node, and it logs one line saying so; retention_boot_holds_held above zero long after start means a group was never hosted again (its restarts exhausted, or it is quarantined), and compaction on that store frees nothing until it is.
Memory admission
Heavy statements take a memory grant before they run and queue in arrival order when the pool is full. Nothing is ever refused, so a rising queue means work is waiting, not failing.
| Metric | Type | Meaning |
|---|---|---|
scramdb_admission_admitted_total | counter | Heavy statements that took a grant. |
scramdb_admission_queued_total | counter | Heavy statements that had to wait for one, counted once each when they first park. |
scramdb_admission_parked_waiters | gauge | Statements parked waiting for a grant right now. Emitted only once a real measurement exists. |
scramdb_admission_outstanding_grant_bytes | gauge | Bytes currently granted and not yet handed back. |
scramdb_admission_waited_past_deadline_total | counter | Statements that waited longer than the warning threshold. Nothing is refused on that threshold; it exists so a long wait is visible. |
scramdb_admission_cancelled_while_queued_total | counter | Statements whose own cancellation fired while they were still waiting. |
scramdb_maintenance_memory_waits_total | counter | CREATE INDEX builds (vector indexes and their REINDEX included) and ANALYZE runs whose working memory (storage.memory.maintenance_bytes) did not fit the pool when asked, and that waited for it instead of failing. The automatic analysis a COPY or bulk INSERT asks for counts here too, and runs once the pool has room. |
scramdb_analyze_on_load_waits_total | counter | Automatic analyses of a COPY or bulk INSERT ([statistics] analyze_on_load) that waited for the load's transaction to commit, so they read the loaded rows. Growth is normal: a load inside a transaction block counts once while it waits for its COMMIT. |
scramdb_admission_bypassed_total | counter | Statements served without consulting the ledger, either because they were light enough to stream or because the gate is switched off. |
scramdb_admission_inherited_nested_total | counter | Nested statements served inside their parent's grant instead of taking a second one. Growth is normal on a workload full of CTEs and routines. |
scramdb_admission_released_at_seam_total | counter | Grants handed back cleanly at statement completion. On ordinary workloads this should climb in step with admitted_total; a gap that keeps growing means grants are being returned late, and is the first thing to check if concurrency degrades over a long run. |
scramdb_admission_measured_grants_total | counter | Statements admitted with a grant sized from their measured memory. |
scramdb_admission_measured_growths_total | counter | Measured grants that had to grow while their statement ran. |
scramdb_admission_heavy_concurrency | gauge | Heavy statements the gate currently admits at once. Live-settable at runtime, independent of the storage.memory.heavy_concurrency config key's value at startup. |
Allocator accounting
Every byte ScramDB's own allocator hands out is counted exactly, the instant it is allocated; see Memory for how the figures below feed the budgets on that page.
| Metric | Type | Meaning |
|---|---|---|
scramdb_heap_live_bytes | gauge | The exact process heap total: every byte the allocator has handed out and not yet freed. |
scramdb_heap_reserve_bytes | gauge | Bytes the allocator has committed but not yet handed to a live allocation (its own retention). |
scramdb_heap_overflow_cells | gauge | Threads currently sharing the heap accounting overflow cell instead of their own. |
scramdb_heap_context_overflows_total | counter | Charge contexts created after the context table was full, sharing a class's permanent slot. |
scramdb_heap_class_bytes | gauge | Live bytes per class (statement, index_build, vector_index, background, cache, cluster, unattributed), labeled class. |
scramdb_heap_escaped_bytes | gauge | Bytes still resident under a class whose owning statement, build or job already ended, labeled class. Steady growth names a component that keeps memory after the work that allocated it has finished. |
scramdb_heap_uncharged_bytes | gauge | Bytes live statements hold in their allocator context past what they reserved through the execution-memory budget; the budgets already decide on the larger figure, so this shows how much work runs ahead of its reservations. |
scramdb_execution_pool_allocated_bytes | gauge | Bytes the execution memory pool currently holds against its budget (the larger of what was reserved and what the allocator measured). |
scramdb_execution_pool_budget_bytes | gauge | The execution memory pool's total budget (execution_memory_bytes). |
scramdb_execution_pool_essential_breach_episodes_total | counter | Times state a query could not give back (rows already built, mid-flight) carried the execution memory pool past its budget and its small overshoot allowance. Reported once per crossing, never once per allocation, so a healthy process reads zero or a small, stable count; a climbing count means concurrent heavy queries are routinely exceeding execution_memory_bytes together, and raising it or running fewer of them at once is the fix. |
Joins over the memory budget and buffer reuse (always present)
A hash join whose build side does not fit the execution-memory budget spills and runs in rounds, and several such joins in one probe stage share one round schedule. When the joins' keys are columns of the probe table itself, the probe side is partitioned to disk once and each pass reads only the rows its rounds can match; the query workers reuse their interpreters and join buffers across morsels. A join whose output for one batch of its probe side is larger than one output batch (a cross join, a key the build side holds many times) emits it in bounded batches, each charged to the statement's memory while it is held, and a plain query over such a join streams them to the client as they are made.
| Metric | Type | Meaning |
|---|---|---|
scramdb_join_round_plans_total | counter | Hash join builds that did not fit the execution-memory budget and ran in rounds. |
scramdb_join_round_passes_total | counter | Passes over a probe side made by hash joins running in rounds. |
scramdb_join_round_multi_plan_stages_total | counter | Probe stages that ran several oversized hash joins in rounds together. |
scramdb_join_round_chunked_plans_total | counter | Hash join builds that split one partition into several rounds to stay within the execution-memory budget. |
scramdb_join_round_partitioned_probes_total | counter | Probe stages of hash joins in rounds whose probe side was partitioned to disk once instead of read on every pass. |
scramdb_join_output_steps_total | counter | Join output batches emitted by a probe whose output was larger than one batch. |
scramdb_unnest_output_steps_total | counter | Unnest output batches emitted for arrays whose elements were more than one batch. |
scramdb_join_output_step_bytes | gauge | Bytes of join and unnest output batches the query workers hold right now. |
scramdb_join_output_step_peak_bytes | gauge | Most bytes of join and unnest output batches the query workers held at once. |
scramdb_execution_interpreters_built_total | counter | Interpreters the query workers built to run morsels on the interpreter. |
scramdb_execution_join_selection_grows_total | counter | Times a query worker grew its reused join selection buffer. |
scramdb_execution_agg_rows_folded_in_runs_total | counter | Grouped aggregate rows folded into the previous row's group without a hash lookup, because their key equals the previous row's. |
scramdb_execution_agg_rows_folded_in_runs_compiled_total | counter | The same, for the rows the compiled tier folded. |
scramdb_execution_agg_run_checks_switched_off_total | counter | Grouped aggregate batches whose consecutive-key check switched off for lack of runs, so input not ordered on its keys pays only a short warm-up of comparisons per batch. |
scramdb_execution_call_argument_views_total | counter | Batch arguments handed to scalar functions as zero-copy views instead of copies. |
The build counters stay bounded by the workers and the batch sizes, never by the rows a statement reads: a count that grows with the input points at a buffer that stopped being reused. The folded-row counters grow on input grouped or sorted by its grouping key, where most rows skip the hash table.
Spilling to disk
A query that outgrows its memory budget writes to disk rather than failing. These count that happening.
| Metric | Type | Meaning |
|---|---|---|
scramdb_agg_spill_partition_flushes_total | counter | Aggregate partitions written to disk instead of breaching the memory budget. |
scramdb_agg_spill_bytes_written_total | counter | Bytes written to aggregate spill files. |
scramdb_agg_spill_rows_written_total | counter | Rows written to aggregate spill files. |
scramdb_agg_spill_fragment_partition_flushes_total | counter | Fragment partitions written to disk for the same reason. |
scramdb_agg_spill_fragment_bytes_written_total | counter | Bytes written to fragment spill files. |
scramdb_agg_spill_fragment_rows_written_total | counter | Entries written to fragment spill files. |
scramdb_spill_tape_live_bytes | gauge | Bytes currently resident in every live spill file. Observed, not capped: spill disk usage is deliberately unbounded so a large query completes rather than failing. Alert on it if the volume is small. |
scramdb_fragment_spill_mgr_init_failures_total | counter | Times the fragment spill directory could not be initialized and was retried. Zero on a healthy process; non-zero points at temp_dir permissions or space. |
scramdb_final_agg_spill_mgr_init_failures_total | counter | The same, for the final aggregate spill directory. |
Bulk load
| Metric | Type | Meaning |
|---|---|---|
scramdb_copy_one_pass_chunks_total | counter | Chunks processed by the single-read COPY parser. Zero means that path never engaged, either because it is switched off or because no eligible COPY has run. |
scramdb_copy_one_pass_reparses_total | counter | Chunks that had to be parsed a second time because the parser's guess about where quoting stood at the chunk boundary was wrong. Against chunks_total this is the miss rate; a file with heavy embedded quoting raises it. |
scramdb_copy_one_pass_reparse_bytes_total | counter | Bytes re-read by those second passes, tracked separately so load throughput is never overstated. |
scramdb_copy_staged_bytes | gauge | Bytes of staged COPY rows and key runs held on this node. |
scramdb_copy_stages_open | gauge | COPY stages holding rows that are not committed yet: a node-local COPY's, and on a cluster node also every replica's stage of a cluster COPY and the coordinating node's stage of the COPY's key checks. |
scramdb_copy_staged_rows_total | counter | Rows COPY staged. |
scramdb_copy_stages_attached_total | counter | COPY stages committed into their tables. |
scramdb_copy_stages_dropped_total | counter | COPY stages discarded without a commit. |
scramdb_copy_staging_refusals_total | counter | COPY batches refused because the staging area was full. Each one failed its COPY with the disk-full error, leaving nothing behind. |
scramdb_copy_stages_recovered_total | counter | COPY stages a restart found and freed. |
scramdb_copy_stages_restored_total | counter | Committed COPY stages a restore or a branch rebuilt from the WAL. |
scramdb_copy_attach_seconds | histogram | Time to attach one COPY stage to its table at commit. |
A COPY writes its rows into staged segments outside the table, invisible to every reader, and attaches them to the table in one step when its transaction commits; a rollback or a crash before the commit frees them whole. copy_staged_bytes is the disk those uncommitted rows hold now. A COPY into a cluster table stages its rows on every replica of every shard group it writes and commits them together (see COPY into a distributed table); copy_stages_open above zero while no COPY is running means a stage was left behind, which a group leader collects after [cluster.dilith] copy_stage_idle_timeout (scramdb_cluster_dilith_copy_stages_collected_total).
Import (COPY ... FROM STDIN)
| Metric | Type | Meaning |
|---|---|---|
scramdb_copy_in_started_total | counter | COPY FROM STDIN statements that started receiving rows from a client. |
scramdb_copy_in_completed_total | counter | COPY FROM STDIN statements that loaded their rows. |
scramdb_copy_in_failed_total | counter | COPY FROM STDIN statements that failed, including those the client aborted. |
scramdb_copy_in_client_aborts_total | counter | COPY FROM STDIN statements the client ended with CopyFail. |
scramdb_copy_in_bytes_total | counter | Bytes of COPY FROM STDIN input handed to loads. |
scramdb_copy_in_slices_total | counter | Slices of COPY FROM STDIN input handed to loads. |
scramdb_copy_in_client_waits_total | counter | Times a COPY FROM STDIN's input waited for its load to take the slices before it. |
An import queues at most [execution] copy_pipeline_depth slices of input for its load and makes the client wait while the queue is full, so copy_in_client_waits_total climbing against copy_in_started_total means the load is slower than the client sends; the server is not buffering the difference.
Export (COPY ... TO STDOUT)
| Metric | Type | Meaning |
|---|---|---|
scramdb_copy_out_streams_total | counter | COPY TO STDOUT statements that started streaming rows to a client. |
scramdb_copy_out_rows_total | counter | Rows COPY TO STDOUT statements streamed to clients. |
scramdb_copy_out_bytes_total | counter | Bytes of COPY TO STDOUT output streamed to clients. |
scramdb_copy_out_failed_total | counter | COPY TO STDOUT statements that ended with an error after they started streaming. |
scramdb_copy_out_client_waits_total | counter | Times a COPY TO STDOUT waited for its client to take rows before rendering more. |
scramdb_copy_out_seconds | histogram | Time from a COPY TO STDOUT's CopyOutResponse to its CopyDone. |
An export holds [execution] copy_pipeline_depth buffers and waits for the client when they are full, so copy_out_client_waits_total climbing against copy_out_streams_total means clients read slower than the server renders; the export is not buffering the difference.
Foreign keys
| Metric | Type | Meaning |
|---|---|---|
scramdb_fk_referencing_keys_checked_total | counter | Referencing keys checked against the table they reference. |
scramdb_fk_referenced_keys_checked_total | counter | Referenced keys that left or changed and were checked for rows still naming them. |
scramdb_fk_violations_total | counter | Statements and commits refused because a foreign key was violated. |
scramdb_fk_actions_total | counter | Referential actions (CASCADE, SET NULL, SET DEFAULT) run on referencing tables. |
scramdb_fk_deferred_checks_total | counter | Foreign key checks deferred to the transaction's commit point. |
scramdb_fk_key_shares_total | counter | Referenced keys a referencing check held shared while it ran, so a concurrent delete of the key waits for it. |
scramdb_fk_key_share_waits_total | counter | Referencing checks that waited for a transaction removing the key they reference. |
scramdb_fk_leaving_key_locks_total | counter | Referenced keys held exclusively because they left or changed. |
scramdb_fk_removed_after_snapshot_total | counter | Referencing checks refused with 40001 because a key their snapshot held was removed by a later commit. |
scramdb_fk_probe_seconds | histogram | Time one foreign key probe statement took. |
Deep diagnostics
These are emitted but are not part of the supported surface, and their names and meanings can change between releases. They are here so a number you see on the endpoint is never unexplained; do not build alerts on them.
| Metric | Type | Meaning |
|---|---|---|
scramdb_header_sidecar_loads_total | counter | Segments whose page headers were loaded entirely from their persisted sidecar file: one sequential read instead of one per page. |
scramdb_header_sidecar_partial_loads_total | counter | Sidecars that covered only part of what was asked for, the ordinary case for a table still growing. |
scramdb_header_sidecar_fallbacks_total | counter | Sidecar reads that fell back to walking page headers because the file was absent, stale or corrupt. |
scramdb_header_sidecar_writes_total | counter | Header sidecars written. |
scramdb_header_sidecar_write_errors_total | counter | Header sidecar writes that failed. The segment is walked next time instead. |
scramdb_header_prefix_pages_walked_total | counter | Pages read by the per-page header walk. A warm store should stop moving it. |
scramdb_page_zero_bytes_full_total | counter | Bytes zeroed over a whole page buffer on release. |
scramdb_page_zero_bytes_partial_total | counter | Bytes zeroed over only the recorded written region. On a warm selective scan this should dominate the full counter. |
scramdb_scan_prune_nanos_total | counter | Core-nanoseconds spent choosing which pages to read. Summed across workers, so it exceeds wall-clock time on a parallel scan. |
scramdb_scan_acquire_nanos_total | counter | Core-nanoseconds spent acquiring pages, pinning them when resident and reading them when not. |
scramdb_scan_decode_nanos_total | counter | Core-nanoseconds spent turning page bytes into chunks. |
scramdb_agg_fold_rows_total | counter | Partial aggregate rows consumed by the final fold. |
scramdb_agg_fold_chunks_parallel_total | counter | Fold chunks large enough to spread across cores. |
scramdb_agg_fold_chunks_in_place_total | counter | Fold chunks small enough to fold on one core. |
scramdb_agg_fold_distinct_values_offered_total | counter | DISTINCT values offered to a fold's deduplication set. |
Licensing metrics
| Metric | Type | Meaning |
|---|---|---|
scramdb_license_valid | gauge | 1 if the active license is valid. |
scramdb_license_edition | gauge/label | Active license edition. |
scramdb_license_node_count | gauge | Nodes currently counted against the license. |
scramdb_license_max_nodes | gauge | Node limit the license allows. |
scramdb_license_read_only_active | gauge | 1 if the license has forced the deployment read-only. |
scramdb_license_days_of_term_remaining | gauge | Emitted only for a term-limited license. |
scramdb_license_data_bytes | gauge | Emitted only for a license with a data-size cap. |
scramdb_license_data_cap | gauge | Emitted only for a license with a data-size cap. |
Cluster mode
Running as a cluster node, additional scramdb_cluster_* metrics are appended to the same /metrics response after the single-node metrics above. Every failure mode below has a counter: a stalled compaction, a quarantined group, a dropped partition frame. Alerting on a ScramDB cluster means reading a number, never grepping logs for a string match.
Unlike the WAL archive metrics above, every scramdb_cluster_* metric is always present once [cluster] is configured, whether or not anything interesting has happened yet. A fresh cluster reports every counter at 0 and every gauge at its idle value; there's no "not measured yet" gap to account for in this family.
Peer transport
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_peer_connects_total | counter | Peer connections established over the cluster transport. |
scramdb_cluster_peer_disconnects_total | counter | Peer connections closed. |
scramdb_cluster_reconnect_attempts_total | counter | Reconnect attempts after a lost peer connection. |
scramdb_cluster_frames_sent_total | counter | Frames sent to peers. |
scramdb_cluster_frames_received_total | counter | Frames received from peers. |
scramdb_cluster_bytes_sent_total | counter | Bytes sent to peers. |
scramdb_cluster_bytes_received_total | counter | Bytes received from peers. |
scramdb_cluster_send_queue_bytes | gauge | Outbound send-queue occupancy across all peer connections, right now. Bounded by send_queue_bytes in [cluster] config (default 64MB) per connection; a value pinned near that bound is backpressure, not idle capacity. |
scramdb_cluster_partition_drops_total | counter | Frames silently dropped by an injected network partition, outbound and inbound combined. Zero on every normal deployment; nonzero means a chaos/test partition is, or was, active. |
scramdb_cluster_fragment_boot_refusals_total | counter | Fragment requests from other nodes answered with a refusal because this node was still starting. |
scramdb_cluster_forward_boot_refusals_total | counter | Requests other nodes forwarded to this node for the metadata group (a DDL, a lock, a node number claim, a member removal) answered as not made because this node was still starting: the sender asks the next node at once instead of waiting out its forward timeout. |
scramdb_cluster_handshake_latency_seconds | histogram | Peer handshake latency. 11 buckets from 1ms to 5s (le="0.001" through le="5.000") plus +Inf. |
scramdb_cluster_connections_closed_total | counter | Peer connections closed, by cause: ping_deadline (no byte from the peer within its dead-peer deadline, the stalled-process case), tcp_timeout (the kernel's user timeout or keepalive gave up), peer_closed, local (shutdown or a duplicate connection), io_error, protocol_error. |
scramdb_cluster_frames_dropped_on_close_total | counter | Frames still queued on a peer connection when it closed, dropped and reported to the subsystems that sent them. |
scramdb_cluster_peer_events_total | counter | Peer connection events published to the cluster's subsystems, by kind: connected and lost. |
scramdb_cluster_peer_event_lag_total | counter | Times a subsystem fell behind the connection event channel and treated every peer as lost once. Zero in normal operation. |
scramdb_cluster_liveness_pings_sent_total | counter | Liveness pings sent on peer connections that had sent nothing for ping_interval. |
scramdb_cluster_liveness_ping_rtt_seconds | histogram | Round trip of a liveness ping, from its send to its answer. 18 buckets from 10µs to 10s plus +Inf. |
scramdb_cluster_dead_peer_deadline_connections | gauge | Open peer connections by how their dead-peer deadline is set, source: floor (the adaptive deadline at dead_peer_timeout_min), measured (lifted by a long round trip), ceiling (held at dead_peer_timeout_max), pinned (a fixed dead_peer_timeout). |
scramdb_cluster_dead_peer_deadline_seconds | histogram | Every dead-peer deadline put in force, at a connection's start and whenever its measured round trip moves it. |
scramdb_cluster_socket_option_failures_total | counter | Socket options (TCP_NODELAY, keepalive, user timeout) the kernel refused on a peer connection. Zero on Linux; nonzero means dead-peer detection relies on the ping deadline alone. |
scramdb_cluster_peers_remembered | gauge | Peers whose connection closed within the grace period and whose handshake details are kept. |
scramdb_cluster_peers_remembered_expired_total | counter | Remembered peers dropped after staying disconnected past the grace period. |
scramdb_cluster_peers_left_total | counter | Peers forgotten because they left the cluster membership. |
scramdb_cluster_dial_backs | gauge | Peers this node is dialing back because this node must open their connection. |
scramdb_cluster_dial_backs_started_total | counter | Times this node started dialing a peer back. |
scramdb_cluster_dial_backs_ended_total | counter | Times this node stopped dialing a peer back. |
scramdb_cluster_peer_identity_refusals_total | counter | Peer connections refused because the certificate does not carry the node name the peer announced. Nonzero means a node's certificate is missing its name; see TLS between cluster nodes. |
scramdb_cluster_duplicate_connections_refused_total | counter | Peer connections refused because this node already accepted one from the same peer on the same lane. A peer that opens a lane twice at once (two seed addresses for one node, or a reconnect racing the connection it replaces) keeps the first and drops the second, so a few are normal; a steady climb means a peer keeps reopening a lane it already holds. |
scramdb_cluster_transport_threads | gauge | Threads the cluster transport runs for its listeners, dials, handshakes and peer connections. |
scramdb_cluster_transport_thread_start_failures_total | counter | Threads the cluster transport could not start. |
During an injected partition, partition_drops_total also counts the dropped liveness pings and every message that was already queued to a cut peer when the partition started (a message already partly sent is finished once the partition heals), and the connection to each cut peer is closed by the ping deadline and re-established every couple of seconds until the partition heals, so peer_connects_total, peer_disconnects_total and reconnect_attempts_total climb with it.
Traffic lanes
Every pair of nodes talks over three connections, one per traffic lane, each on its own port: control (membership, consensus, the commit protocol, forwarded DDL and lock calls, shuffle fetches), interactive (replies to forwarded point reads) and bulk (query data exchange, fragment dispatch with its broadcast join side, snapshot transfer, the rows a cluster COPY stages). A message larger than [cluster.transport] chunk_size travels as chunks, and a lane's connection takes its streams in turn a chunk at a time. Every stream has its own window of its lane, so a subsystem that stops reading holds only its window and stops only its own senders.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_lane_bytes_sent_total | counter | Bytes written to peers, by lane. |
scramdb_cluster_lane_frames_sent_total | counter | Frames written to peers, by lane. |
scramdb_cluster_lane_bytes_received_total | counter | Bytes read from peers, by lane. |
scramdb_cluster_lane_frames_received_total | counter | Frames read from peers, by lane. |
scramdb_cluster_lane_chunks_sent_total | counter | Chunks of large messages written to peers, by lane. |
scramdb_cluster_lane_chunks_received_total | counter | Chunks of large messages read from peers, by lane. |
scramdb_cluster_lane_messages_reassembled_total | counter | Large messages reassembled from chunks, by lane. |
scramdb_cluster_lane_credit_frames_sent_total | counter | Frames that returned stream credit to a peer, by lane. |
scramdb_cluster_lane_connections | gauge | Open peer connections, by lane. |
scramdb_cluster_lane_older_protocol_connections | gauge | Open peer connections that speak an older cluster protocol version, by traffic lane. |
scramdb_cluster_lane_protocol_fallbacks_total | counter | Dials made again with the older cluster protocol version a peer asked for, by traffic lane. |
scramdb_cluster_lane_stream_window_bytes | gauge | Receive window granted to each stream on the newest connection, by lane: the lane's size shared by the peers then connected. |
scramdb_cluster_lane_queued_bytes | gauge | Bytes held by each traffic queue, by lane and direction: send (queued to send) or receive (received and not yet consumed). |
scramdb_cluster_lane_capacity_bytes | gauge | Size of each traffic queue, by lane and direction. |
scramdb_cluster_lane_backpressure_total | counter | Sends refused because a traffic queue or a stream window was full, by lane and stream (membership, metadata_consensus, shard_consensus, commit_protocol, ddl_forward, lock_forward, shuffle_fetch, point_read, exchange, fragment_dispatch, snapshot_transfer, copy_staging, other). |
scramdb_cluster_lane_drain_waits_total | counter | Sends that waited for a full traffic queue to drain instead of failing. |
scramdb_cluster_lane_drain_wait_seconds | histogram | Time a send waited for a full traffic queue to drain. |
scramdb_cluster_lane_connection_threads | gauge | Reader and writer threads of peer connections, by lane: two per open connection. |
backpressure_total on the bulk lane is flow control doing its job under a large query. On the control lane it should stay at zero: a nonzero shard_consensus or metadata_consensus count means consensus traffic found its queue full, and each such event is also logged once.
Raft group health and recovery
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_group_quarantines_total | counter | Raft group storage directories quarantined: at boot for deeper corruption (a bad header, a mid-prefix CRC failure, an undecodable hard state), or at runtime when a group stopped for a reason a restart cannot fix. Does not count a benign torn-tail truncation; that's a row below. |
scramdb_cluster_quarantined_groups | gauge | Groups currently quarantined, awaiting a membership-replace rejoin. This is the line to alert on. |
scramdb_cluster_group_rejoins_total | counter | Quarantined groups that completed their membership-replace rejoin. |
scramdb_cluster_term_changes_total | counter | Metadata group (group0) raft term changes this node has observed: an election, a term bump from a higher-term message, or a step-down. Node-local, not a cluster-wide total: each node counts its own view of the same election. |
scramdb_cluster_group_actor_restarts_total | counter | Shard groups this node restarted in place after a transient stop, within a bounded restart budget; one that keeps stopping is quarantined instead. |
scramdb_cluster_torn_tail_truncations_total | counter | Benign torn-tail truncations at raft log open: the normal crash-recovery path, made visible instead of silent. |
scramdb_cluster_quarantined_groups is restart-honest. It is never incremented or decremented as events happen; every host-reconcile round recomputes it from scratch, from the actual set of groups gated right now, and overwrites the gauge with that count. A crash that loses an in-flight update, or a process restart mid-recovery, can't leave this gauge stuck on a stale nonzero reading or silently reset to a wrong zero: the very next reconcile round always reflects what's really quarantined at that moment, never a running total of past events.
Consensus threads
Every replication group and the metadata group of a node run on one pool of consensus threads ([cluster.consensus] threads), never a thread per group. A group runs one step at a time and hands its thread back, so a slow group holds one thread and never another group.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_consensus_pool_threads | gauge | Threads of the pool that drives the consensus groups. |
scramdb_cluster_consensus_pool_queued | gauge | Consensus groups and cluster tasks waiting for a consensus thread. Near zero on a healthy node; a value that stays high means the threads cannot keep up. |
scramdb_cluster_consensus_pool_panics_total | counter | Consensus groups and cluster tasks stopped by a panic. Zero on a healthy node; the log names the cause. |
scramdb_cluster_consensus_mailbox_refused_total | counter | Messages a consensus group refused because its queue was full or closed. |
Exchange threads
The distributed query work other nodes send a node (their fragments, the rows arriving, and the results sent back) runs on its own pool of exchange threads ([cluster.exchange] threads), never on the consensus threads and never on the threads that run queries.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_exchange_pool_threads | gauge | Threads of the pool that runs distributed query work sent by other nodes. |
scramdb_cluster_exchange_pool_queued | gauge | Distributed query tasks waiting for an exchange thread. Near zero on a healthy node; a value that stays high means the threads cannot keep up. |
scramdb_cluster_exchange_pool_panics_total | counter | Distributed query tasks stopped by a panic. Zero on a healthy node; the log names the cause. |
Cluster services
Requests other nodes send this node's cluster services (lock and DDL forwarding, snapshot transfer, query fragments and shuffle fetches) are served on engine threads, each service with its own bound on the requests it serves at once.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_stream_requests_total | counter | Requests from other nodes this node's cluster services started serving. |
scramdb_cluster_stream_requests_in_flight | gauge | Requests from other nodes this node's cluster services are serving now. |
scramdb_cluster_stream_requests_waited_total | counter | Requests that waited for a free place because their service was at its bound. |
scramdb_cluster_stream_frames_refused_total | counter | Frames dropped because their service had stopped or its queue was full. |
Cluster-wide LOCK TABLE grants and transaction advisory locks are released at the end of the transaction or the session that holds them, a session that ends with its transaction open included:
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_lock_grants_released_total | counter | Cluster-wide lock grants released at the end of a transaction or session. |
scramdb_cluster_lock_grant_release_failures_total | counter | Cluster-wide lock grants whose release at the end of a transaction or session failed. |
scramdb_cluster_lock_detached_releases_total | counter | Sessions that ended holding cluster-wide lock grants, released after they closed. |
Group commit and durability
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_group_commit_ops_total | counter | Raft ops served, summed across every group-commit wake-up. Divide by group_commit_wakeups_total for the achieved batch size. A lone proposal (nothing to batch with) contributes 1 op and 0 wake-ups, so an idle system reports no batching instead of a fake 1.0. |
scramdb_cluster_group_commit_wakeups_total | counter | Wake-ups that served more than one op, i.e. wake-ups that actually batched. |
scramdb_cluster_group_commit_max_batch | gauge | The largest batch any wake-up has served so far, a high-water mark. Pinned at your configured cap says "raise the cap"; comfortably under it says the batching window is already draining faster than proposals arrive. |
scramdb_cluster_durability_flushes_total | counter | Durability flushes, one per fsync. Divide into group_commit_ops_total for fsyncs-per-commit, the number batching exists to lower. |
scramdb_cluster_log_compactions_total | counter | Raft log compactions performed. |
scramdb_cluster_log_entries_compacted_total | counter | Raft log entries reclaimed by compaction, summed. |
Zero log_compactions_total is normal on a cluster whose log never reaches the compaction trigger. A flat zero on a cluster that's otherwise busy (group_commit_ops_total climbing) is not normal: compaction is waiting, either for a state snapshot that keeps failing (state_snapshot_failures_total or state_snapshots_abandoned_total climbing: a full disk, or a [cluster.log_compaction] memory_bytes budget too small for the group's state) or for a replica that is behind but still inside the lag limit. A replica lagging further than replica_lag_limit_entries no longer holds the log: when it returns it is repaired from a state snapshot, so a dead replica bounds the log instead of growing it forever. The waiting case also logs on its own, but the metric is what you'd wire an alert to; see Alerting on cluster metrics below.
Shared raft log
Every shard group a node hosts writes its replication log into the node's one shared raft log: one writer makes a whole batch of appends from every group durable with one data sync, so a transaction that touches several groups costs the node one sync, not one per group. Its segment files are reused a whole segment at a time once nothing in them is needed; [cluster.consensus] log_segment_size sets their size (by default a share of the data volume). The metadata group keeps its own log.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_shared_log_drains_total | counter | Batches the writer of the node's shared raft log wrote, each ending in one sync. |
scramdb_cluster_shared_log_syncs_total | counter | Data syncs of the node's shared raft log. |
scramdb_cluster_shared_log_group_syncs_total | counter | Groups whose writes a sync of the node's shared raft log covered, summed over its syncs. Divided by syncs_total, the mean groups one sync made durable. |
scramdb_cluster_shared_log_records_total | counter | Records written to the node's shared raft log. |
scramdb_cluster_shared_log_bytes_written_total | counter | Bytes written to the node's shared raft log. |
scramdb_cluster_shared_log_copied_bytes_total | counter | Entry payload bytes copied into write buffers of the node's shared raft log; larger payloads are written in place. |
scramdb_cluster_shared_log_truncations_total | counter | Truncate records written for a replica's conflicting log suffix. |
scramdb_cluster_shared_log_segments_created_total | counter | Segment files the node's shared raft log started. |
scramdb_cluster_shared_log_segments_removed_total | counter | Segment files of the node's shared raft log removed once nothing in them was needed. |
scramdb_cluster_shared_log_compactions_requested_total | counter | Groups asked to compact their log because they held the oldest segment of the node's shared raft log. |
scramdb_cluster_shared_log_relocated_records_total | counter | Records rewritten at the head of the node's shared raft log so an old segment could be removed. |
scramdb_cluster_shared_log_relocated_bytes_total | counter | Bytes rewritten at the head of the node's shared raft log so an old segment could be removed. |
scramdb_cluster_shared_log_sync_failures_total | counter | Failed writes or syncs of the node's shared raft log; the first one stops every group on the node. |
scramdb_cluster_shared_log_torn_bytes_total | counter | Bytes of an unfinished write cut from the end of the node's shared raft log at start. |
scramdb_cluster_shared_log_migrated_groups_total | counter | Groups whose own raft log was moved into the node's shared raft log. |
scramdb_cluster_shared_log_segments | gauge | Segment files of the node's shared raft log. |
scramdb_cluster_shared_log_bytes | gauge | Bytes of the node's shared raft log on disk. |
scramdb_cluster_shared_log_live_bytes | gauge | Bytes of the node's shared raft log still needed by a group. |
scramdb_cluster_shared_log_groups | gauge | Groups with records in the node's shared raft log or a store open on it. |
scramdb_cluster_shared_log_stopped | gauge | 1 once a failed write or sync stopped the node's shared raft log, else 0. |
scramdb_cluster_shared_log_segment_bytes | gauge | The segment size of the node's shared raft log in force. |
scramdb_cluster_shared_log_scan_milliseconds | gauge | How long the node's shared raft log took to read its segments at start. |
scramdb_cluster_shared_log_sync_seconds | histogram | Time one data sync of the node's shared raft log took. |
scramdb_cluster_shared_log_drain_seconds | histogram | Time from the start of a batch's write to its sync returning. |
shared_log_stopped at 1 is the one to page on: a write or sync of the log failed, every shard group on the node stopped rather than trust a disk whose state is no longer known, and the node must be restarted (nothing is retried). shared_log_bytes far above shared_log_live_bytes while compactions_requested_total climbs means an idle group holds the oldest segment; the log asks it to compact and then moves its few live records forward, which relocated_bytes_total counts.
State snapshots and log compaction
A data group compacts its replication log behind a snapshot of its commit state: the open transactions' locks and staged writes, the recent versions of every row it wrote, and its read and commit watermarks, taken at one applied log position. A restart restores the newest snapshot and replays only the log above it; a replica that fell behind the compaction point installs the leader's snapshot and rebuilds the rows it missed from it instead of copying the whole group. The cluster metadata log compacts the same way. When a group compacts, how far a replica may lag and how much memory snapshots may hold are set under [cluster.log_compaction]; each count adapts to the group's own snapshot and entry sizes unless you pin it.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_state_snapshots_taken_total | counter | State snapshots taken for a log compaction. |
scramdb_cluster_state_snapshot_failures_total | counter | State snapshots that could not be taken (an I/O error, a full disk). |
scramdb_cluster_state_snapshots_abandoned_total | counter | State snapshots given up to stay within their memory budget. |
scramdb_cluster_state_snapshot_last_bytes | gauge | Size of the most recent state snapshot taken on this node. |
scramdb_cluster_state_snapshot_bytes_total | counter | Bytes of every state snapshot taken. |
scramdb_cluster_state_snapshot_duration_seconds | histogram | Time from the start of a state snapshot to its completion. A snapshot is written in small steps between applied batches, so the group keeps applying while it runs. |
scramdb_cluster_state_snapshot_step_seconds | histogram | Time one of those steps held a group's apply. |
scramdb_cluster_state_snapshot_memory_bytes | gauge | Memory held right now by state snapshots being taken, sent or received. |
scramdb_cluster_state_snapshot_memory_limit_bytes | gauge | Memory they may hold at once (memory_bytes, or the derived default). |
scramdb_cluster_state_snapshot_restores_total | counter | Groups restored from a state snapshot at startup. |
scramdb_cluster_state_snapshot_missing_total | counter | Groups found at startup with a compacted log and no state snapshot to restore from. Zero in a healthy deployment; anything else means that replica needs a full copy. |
scramdb_cluster_state_snapshot_restore_seconds | histogram | Time to restore one group's state at startup. |
scramdb_cluster_log_entries_replayed_total | counter | Log entries replayed at startup on top of the restored state. |
scramdb_cluster_log_torn_tail_replays_total | counter | Groups found at startup with their log cut below what they had applied (a crash, or a disk that lost part of the log). The group takes the cut entries from its peers again and replays them without applying their rows a second time. |
scramdb_cluster_log_compactions_awaiting_snapshot_total | counter | Log compactions held back until a state snapshot covers the open transactions. |
scramdb_cluster_log_compactions_past_lagging_replica_total | counter | Log compactions that passed a replica lagging beyond the lag limit. That replica is repaired from the snapshot when it returns. |
scramdb_cluster_state_snapshot_installs_total | counter | State snapshots from a leader installed on this node. |
scramdb_cluster_state_snapshot_installs_incomplete_total | counter | Snapshot installs whose skipped rows come from a full copy (every install: a snapshot carries no row). |
scramdb_cluster_state_snapshot_fetches_total | counter | State snapshots fetched from a peer, in bounded chunks. |
scramdb_cluster_state_snapshot_fetch_failures_total | counter | State snapshot fetches from a peer that failed; the fetch retries with a growing pause. |
scramdb_cluster_state_snapshot_fetched_bytes_total | counter | Bytes of state snapshots fetched from peers. |
scramdb_cluster_state_snapshot_chunks_served_total | counter | State snapshot chunks served to peers. |
scramdb_cluster_replica_copies_started_total | counter | Full copies of a shard group's rows this node started after a snapshot install. |
scramdb_cluster_replica_copies_completed_total | counter | Full copies of a shard group's rows this node completed. |
scramdb_cluster_replicas_awaiting_rows | gauge | Shard group replicas on this node waiting for a full copy; they serve no read. A value that stays above zero means a copy cannot finish: check replica_copy_attempt_failures_total here and replica_copy_build_failures_total on the group's voters. |
scramdb_cluster_replica_copy_attempt_failures_total | counter | Attempts to copy a shard group's rows from a peer that failed and were retried. |
scramdb_cluster_replica_copy_rows_total | counter | Rows this node received in full copies of shard groups. |
scramdb_cluster_replica_copy_bytes_total | counter | Bytes this node received in full copies of shard groups. |
scramdb_cluster_replica_copy_seconds | histogram | Time from a snapshot install to its full copy being complete. |
scramdb_cluster_replica_copies_built_total | counter | Full copies of a shard group's rows this node built for a peer. |
scramdb_cluster_replica_copy_build_failures_total | counter | Full copies of a shard group's rows this node could not build for a peer. |
scramdb_cluster_replica_copy_chunks_served_total | counter | Chunks of full copies this node served to peers. |
scramdb_cluster_metadata_log_compactions_total | counter | Cluster metadata log compactions. |
scramdb_cluster_metadata_state_restores_total | counter | Cluster metadata states restored or installed from a state snapshot. |
scramdb_cluster_log_compaction_trigger_entries | gauge | Retained log entries at which a data group last decided to compact. |
scramdb_cluster_replica_lag_limit_entries | gauge | Log entries a replica may lag before it stops holding the leader's log. |
scramdb_cluster_log_compaction_trigger_source | gauge | Why the trigger has its value: 0 configured, 1 the floor (nothing measured yet, or measured below it), 2 measured from the group's snapshot and entry sizes, 3 the ceiling. |
Cluster archive
With [storage.wal.archive] enabled, each shard group's leader ships the group's committed log, and the metadata group's leader the metadata log, to the archive destination, for cluster restore. Each segment lands under a key only one leader can create, so a leader that lost its leadership can never overwrite what the new one shipped, and once it finds the group archived in a newer term it stops archiving that group.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_archive_segments_total | counter | Log segments this node shipped to the cluster archive. |
scramdb_cluster_archive_entries_total | counter | Log entries this node shipped to the cluster archive. |
scramdb_cluster_archive_bytes_total | counter | Bytes of log segments this node shipped to the cluster archive. |
scramdb_cluster_archive_fenced_total | counter | Log segments refused because another leader had already archived that log position. |
scramdb_cluster_archive_superseded_total | counter | Groups this node stopped archiving because a leader of a newer term archives them. |
scramdb_cluster_archive_bases_total | counter | Group bases (a state image and a full copy of the rows) this node shipped. |
scramdb_cluster_archive_images_total | counter | Group state images this node shipped alone, between bases. |
scramdb_cluster_archive_gaps_total | counter | Times a group's log was compacted past the archive and no base could be shipped. |
scramdb_cluster_archive_errors_total | counter | Failed attempts to write to the cluster archive. |
scramdb_cluster_archive_pruned_total | counter | Archive objects the retention pass removed. |
scramdb_cluster_archive_time_samples_total | counter | Time index samples this node shipped to the cluster archive. |
scramdb_cluster_archive_lag_entries | gauge | Committed log entries of the groups this node leads that are not archived yet. |
scramdb_cluster_archive_groups | gauge | Groups whose log this node ships now. |
scramdb_cluster_archive_ship_seconds | histogram | Time to ship one log segment to the cluster archive. |
A few fenced_total around a leader change are expected: the old leader's last attempt is refused because the new leader already archived that position. gaps_total above zero means a group's history has a hole a restore cannot cross (the group compacted its log past what was archived and no base could be shipped): a restore to a point inside that hole is refused, and points after the group's next base restore again. lag_entries growing without bound means the destination cannot keep up.
Distributed transaction commit
What the commit protocol of distributed transactions ([cluster.dilith] in the config file) did on this node.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_dilith_commits_total | counter | Transactions committed. |
scramdb_cluster_dilith_aborts_total | counter | Transaction attempts aborted, by reason: conflict (another transaction held a pending write or a read fence the vote needed), stale_read (a version above a read position, or a read below the retention floor), outdated_vote (older than the group's memory of decisions), wait_die (the younger of two transactions on one row), recovery (a recovery found a group that never voted). |
scramdb_cluster_dilith_commit_latency_seconds | histogram | Time from a commit request to its answer. |
scramdb_cluster_dilith_vote_round_seconds | histogram | Time from sending a transaction's votes to its last vote answer. |
scramdb_cluster_dilith_votes_appended_total | counter | Vote entries appended by group leaders. |
scramdb_cluster_dilith_precheck_refusals_total | counter | Votes a group leader refused before appending anything. |
scramdb_cluster_dilith_parks_total | counter | Requests parked at a group behind an undecided transaction. |
scramdb_cluster_dilith_write_free_certified_total | counter | Transactions that wrote nothing certified without a log entry. |
scramdb_cluster_dilith_fences_appended_total | counter | Fence entries appended for analytical cuts. |
scramdb_cluster_dilith_fences_avoided_total | counter | Analytical cuts a group already covered without a fence entry. |
scramdb_cluster_dilith_recoveries_started_total | counter | Recoveries started for undecided transactions. |
scramdb_cluster_dilith_recoveries_decided_total | counter | Recoveries that reached a verdict. |
scramdb_cluster_dilith_recoveries_behind_waiter_total | counter | Recoveries started by a request that waited too long behind an undecided transaction. |
scramdb_cluster_dilith_recoveries_abandoned_vote_total | counter | Recoveries started for votes pending too long with nobody waiting. |
scramdb_cluster_dilith_recoveries_handed_over_total | counter | Recoveries handed to the group leader that holds the pending vote. |
scramdb_cluster_dilith_recoveries_from_record_total | counter | Recoveries decided from the leader's own record of the transaction. |
scramdb_cluster_dilith_status_probes_total | counter | Status queries for transactions whose outcome a client did not receive. |
scramdb_cluster_dilith_sweeps_total | counter | Sweep entries appended to decide transactions that never voted at a group. |
scramdb_cluster_dilith_sweeps_refused_total | counter | Sweeps refused because they were older than the group's memory of decisions. |
scramdb_cluster_dilith_sweeps_resent_total | counter | Sweeps sent again with a fresher decision count. |
scramdb_cluster_dilith_votes_resent_total | counter | Vote requests sent again to groups that had not answered, one per group. |
scramdb_cluster_dilith_outcomes_resent_total | counter | Transaction outcomes sent again to groups that had not acknowledged them, one per group. |
scramdb_cluster_dilith_aborts_without_record_total | counter | Abort outcomes acknowledged without recording anything at a group holding no record of the transaction. |
scramdb_cluster_dilith_decisions_forgotten_total | counter | Decisions forgotten once outside the decision window. |
scramdb_cluster_dilith_staged_vote_rounds_total | counter | Vote rounds that asked the groups likely to refuse before the others. |
scramdb_cluster_dilith_copy_stage_entries_total | counter | COPY stage entries applied on this node's replicas. |
scramdb_cluster_dilith_copy_stages_dropped_total | counter | COPY stages dropped on this node's replicas without a commit. |
scramdb_cluster_dilith_copy_stage_vote_refusals_total | counter | Votes refused because the COPY rows they name were not all staged or conflicted. |
scramdb_cluster_dilith_copy_stages_collected_total | counter | COPY stages a group leader on this node dropped after their owner stopped renewing them. |
scramdb_cluster_dilith_copy_stages_refused_full_total | counter | COPY stage entries a group leader on this node refused because its staging area had no room. |
scramdb_cluster_dilith_copy_stages_refused_stale_total | counter | COPY stage entries a group leader on this node sent back to be checked again because the table moved on since their rows were checked. |
scramdb_cluster_dilith_owed_commits_collected_total | counter | Commits whose discharge was lost that a group leader on this node collected. |
scramdb_cluster_dilith_table_entries | gauge | Entries each commit protocol table on this node holds, by table. |
scramdb_cluster_dilith_owed_commits | gauge | Commits on this node's replicas not yet discharged. |
scramdb_cluster_dilith_pending_votes | gauge | Votes accepted on this node's replicas and not yet decided. |
scramdb_cluster_dilith_recent_verdicts | gauge | Recent transaction outcomes this node remembers to answer recoveries. |
scramdb_cluster_dilith_discharges_queued | gauge | Discharge notices waiting for their batch to be sent. |
scramdb_cluster_dilith_refusing_groups | gauge | Groups that refused this node's votes recently. |
scramdb_cluster_dilith_state_bytes | gauge | Memory the commit protocol's tables hold on this node, by part: pending_votes (votes accepted and not decided, with their writes and read fences), owed_commits (commits not yet discharged), kept (the decisions and versions the decision and retention windows keep, and COPY stage records), outcomes (outcomes and discharges a group leader holds for its next entry or has in flight), reserved (room reserved for vote appends in flight), waiting (requests parked at a group, and votes waiting for room), scratch (the node's reusable decode and step buffers). |
scramdb_cluster_dilith_state_share_bytes | gauge | Memory the commit protocol's tables may hold on this node ([cluster.dilith] protocol_state_memory). |
scramdb_cluster_dilith_state_peak_bytes | gauge | Most memory the commit protocol's tables and reserved room have held on this node. |
scramdb_cluster_dilith_votes_waiting_for_memory | gauge | Votes waiting on this node for the commit protocol's memory share to have room. |
scramdb_cluster_dilith_memory_parks_total | counter | Votes a group leader on this node held until the commit protocol's memory share had room. |
scramdb_cluster_dilith_memory_past_share_total | counter | Votes a group leader on this node let past a full memory share so the oldest waiting transaction could go on. |
aborts_total{reason="conflict"} and {reason="wait_die"} count contention between transactions on the same rows. A recoveries_started_total that keeps growing on a healthy network means transactions are being left undecided, and is worth a look at the nodes' logs. copy_stages_collected_total rising means COPY statements stop or lose their node while staging; copy_stage_vote_refusals_total rising means a COPY's vote meets concurrent writers of the rows it stages, and the COPY fails with 40001. copy_stages_refused_stale_total counts batches a COPY checked again because its groups moved on while the batch travelled; the COPY goes on. copy_stages_refused_full_total rising means the staging area is too small for the COPYs running (53100). owed_commits_collected_total rising means nodes stop right after answering commits; every such commit is finished by its groups. table_entries shows every table of the commit protocol and what it holds: each one stays bounded by the work in flight.
The commit protocol's tables live in a memory share of their own (state_share_bytes), counted in the process memory budget. A group leader appends a new vote only while the share has room; a vote that finds none waits, oldest transaction first, and is appended as soon as a decision, a discharge or the protocol's windows free room. It is never refused for memory. votes_waiting_for_memory above zero for long, or memory_parks_total climbing, means the share is too small for the transactions in flight: raise [cluster.dilith] protocol_state_memory. memory_past_share_total counts the oldest waiting transaction let through a full share when nothing on the node could free room without it, which is what keeps the cluster moving; its rise says the same.
The commit protocol on this node
The commit protocol runs as one task on the node's consensus threads, one run at a time, and serves every shard group this node holds a replica of.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_dilith_run_seconds | histogram | Time one run of the commit protocol's host took. |
scramdb_cluster_dilith_runs_total | counter | Runs of the commit protocol's host. |
scramdb_cluster_dilith_run_inputs_total | counter | Inputs the commit protocol's host stepped, over all runs. |
scramdb_cluster_dilith_queued | gauge | Inputs waiting for the commit protocol's host, by kind: control, decide, frame, client. |
scramdb_cluster_dilith_step_errors_total | counter | Inputs the commit protocol could not step, by input: apply, message, timer, leadership, append_decided, restored, client, hosting. |
scramdb_cluster_dilith_reply_slots | gauge | Client calls waiting for an answer from the commit protocol. |
scramdb_cluster_dilith_memory_bytes | gauge | Bytes the commit protocol holds on this node, as last measured. |
scramdb_cluster_dilith_appends_in_flight | gauge | Log entries the commit protocol proposed whose decision has not come back. |
scramdb_cluster_dilith_timers | gauge | Timers the commit protocol keeps armed. |
scramdb_cluster_dilith_frames_refused_total | counter | Frames the transport refused to send; the protocol sends again. |
scramdb_cluster_dilith_appends_lost_total | counter | Proposed log entries lost because the group's leader could not take them. |
scramdb_cluster_commits_in_doubt_total | counter | Commits whose outcome was not known when the client stopped waiting. |
scramdb_cluster_dilith_commits_in_flight | gauge | Commits this node's commit protocol drives whose outcome is not decided yet, as last measured. |
scramdb_cluster_dilith_commits_abandoned | gauge | Commits still in flight here whose client stopped waiting for the outcome, as last measured. |
scramdb_cluster_dilith_calls_abandoned_total | counter | Calls to the commit protocol whose caller stopped waiting, forgotten with what waited to answer them. |
scramdb_cluster_dilith_stage_rounds_resent_total | counter | Rounds of staged rows sent again because their answers did not come in time. |
scramdb_cluster_dilith_stage_requests_lost_total | counter | Requests of staged rows asked again at once because their shard group changed leader. |
scramdb_cluster_dilith_host_stopped | gauge | 1 when the commit protocol on this node stopped after a failure. |
scramdb_cluster_dilith_hosted_groups | gauge | Shard groups the commit protocol hosts on this node. |
scramdb_cluster_dilith_groups_hosted_total | counter | Shard groups the commit protocol started hosting on this node. |
scramdb_cluster_dilith_groups_released_total | counter | Shard groups the commit protocol stopped hosting on this node. |
scramdb_cluster_dilith_decide_requests_dropped_total | counter | Shard group requests dropped because the group is not hosted here or started again. |
scramdb_cluster_dilith_node_number | gauge | This node's number in transaction ids, 0 until the cluster gave it one. |
scramdb_cluster_dilith_node_number_claims_total | counter | Requests this node made to the cluster for its node number. |
scramdb_cluster_dilith_node_number_waits_total | counter | Commit attempts that waited for this node's node number instead of starting at once. |
dilith_host_stopped at 1 means the node can no longer commit transactions and must be restarted; every waiting client was answered with an error. A commits_in_doubt_total that grows means clients stopped waiting before their commit's outcome was known: the commit may still have happened. stage_rounds_resent_total grows when a node sending a COPY's rows, or a transaction's rows too large for one commit message, lost a round to a leader change it could not see or to a node that stopped; one that grows while no leader changed means rounds are taking far longer than the rounds before them.
Committed entries through the commit protocol
Each shard group hands its committed log entries to the commit protocol, which decides the rows they install, and hands the protocol's own entries to the group's log.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_dilith_decide_requests_total | counter | Batches of committed entries the data groups handed to the commit protocol. |
scramdb_cluster_dilith_decide_entries_total | counter | Committed entries the data groups handed to the commit protocol. |
scramdb_cluster_dilith_decided_rows_total | counter | Rows the commit protocol decided for the data groups to install. |
scramdb_cluster_dilith_decide_outstanding | gauge | Requests of the data groups the commit protocol has not answered yet. |
scramdb_cluster_dilith_decide_seconds | histogram | Time from handing committed entries to the commit protocol to receiving their rows. |
scramdb_cluster_dilith_decide_failures_total | counter | Data groups stopped because the commit protocol could not decide an entry of theirs. |
scramdb_cluster_dilith_leadership_reports_total | counter | Leader changes of the data groups reported to the commit protocol. |
scramdb_cluster_dilith_proposal_batches_total | counter | Batches of commit protocol entries handed to a data group's log. |
scramdb_cluster_dilith_proposal_batches_refused_total | counter | Batches of commit protocol entries a data group's log could not take. |
Applying committed writes
A committed entry is applied in two steps on the node's apply threads ([cluster.apply] threads), never on the threads that run consensus, replication and client connections. First it is decided: the commit protocol steps it in memory, in log order (the session that committed it already has its answer, given once every group the transaction touched held a durable yes vote). Then its rows are written, and one data log sync per batch makes a whole batch of them durable before the group's applied position moves. Reads wait for that applied position, so a row is never read before it is applied and synced. On a cluster node the consensus threads keep a share of the cores of their own ([cluster.apply] reserve_io_cores), so heavy queries cannot delay them.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_apply_decide_to_durable_seconds | histogram | Time from deciding a batch of committed entries to its effects being durable: how far the applied position trails the answers sessions already have. |
scramdb_cluster_apply_effect_batches_total | counter | Batches of committed entries whose effects were made durable. |
scramdb_cluster_apply_effect_entries_total | counter | Committed entries whose effects were made durable. Divide by apply_effect_batches_total for the mean batch size. |
scramdb_cluster_apply_data_syncs_total | counter | Data log syncs the apply path requested for those effects. Divide by apply_effect_batches_total: one per batch (zero for a batch with nothing to write) is the design; more means a write path is syncing on its own. |
scramdb_cluster_apply_last_batch_entries | gauge | Entries in the most recent effect batch. |
scramdb_cluster_apply_last_batch_data_syncs | gauge | Data log syncs the most recent effect batch requested. |
scramdb_cluster_apply_decided_batches_waiting | gauge | Decided batches waiting for their effects. Climbing steadily means the disk or the apply threads cannot keep up; each group's backlog stays bounded by its apply mailbox either way. |
scramdb_cluster_apply_stops_total | counter | Groups whose apply stopped on an error. Such a replica is quarantined and rebuilt; see the logs for the reason. |
scramdb_cluster_apply_pool_threads | gauge | Apply threads running. |
scramdb_cluster_apply_pool_queued | gauge | Groups waiting for an apply thread. |
scramdb_cluster_apply_effect_slots_reused_total | counter | Installed rows whose effect was built in a slot an apply thread keeps, its key buffer reused. |
scramdb_cluster_apply_effect_slots_fresh_total | counter | Installed rows whose effect needed a new slot, so its key buffer allocated. Past warm-up it grows only when a batch is larger than any before it on that thread. |
scramdb_cluster_apply_effect_images_decoded_total | counter | Installed rows whose image decode had to allocate memory of its own. |
scramdb_cluster_io_runtime_busy_microseconds_total | counter | Time the consensus and connection threads spent running work. Its rate divided by io_runtime_workers is their utilization. |
scramdb_cluster_io_runtime_workers | gauge | Consensus and connection threads. |
Replication flow control
A leader streams entries to each follower inside a bounded in-flight window (max_inflight_bytes and max_inflight_msgs in [cluster.consensus]), which by default adapts under that ceiling to the bandwidth-delay product measured toward each follower. Each follower's stream is in one of three states: probe (finding where the follower's log ends, one message outstanding), replicate (streaming inside the window) or snapshot (a snapshot is outstanding). The counters and gauges below sum every group this node leads.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_replication_transitions_total | counter | Follower stream state changes, labelled state (probe, replicate, snapshot) and cause: into probe by leader_start, reject (the follower refused an append or a snapshot), connection_loss or stall (two heartbeat checks with entries in flight and no answer); into replicate by acknowledged or snapshot_installed; into snapshot by behind_compaction. |
scramdb_cluster_replication_sends_paused_total | counter | Sends held back because a follower's in-flight window was full. |
scramdb_cluster_replication_append_messages_total | counter | AppendEntries messages sent to followers, heartbeats included. |
scramdb_cluster_replication_append_entries_total | counter | Log entries sent to followers. |
scramdb_cluster_replication_append_bytes_total | counter | Log entry bytes sent to followers. Divided by the bytes proposed and by the follower count, this is the copies of each entry the network carried: about one in steady operation. |
scramdb_cluster_replication_rejects_total | counter | Append and snapshot rejects received from followers. |
scramdb_cluster_replication_snapshot_resends_total | counter | Snapshots sent again to a follower that had not answered within the resend backoff. |
scramdb_cluster_replication_inflight_messages | gauge | Unacknowledged AppendEntries messages to followers, right now. |
scramdb_cluster_replication_inflight_bytes | gauge | Unacknowledged log entry bytes sent to followers, right now. |
scramdb_cluster_replication_window_bytes | gauge | The in-flight window toward a peer, summed over the groups this node leads, labelled peer and reason: ceiling (the configured max_inflight_bytes, or not yet measured), measured (twice the measured bandwidth-delay product), floor (two of the stream's own messages) or pinned (adaptive_inflight_window = false). |
scramdb_cluster_replication_round_trip_milliseconds | gauge | The smoothed replication round trip to a peer, from sending an AppendEntries to its acknowledgement (the follower's fsync included), as last measured. |
scramdb_cluster_replication_entries_committed_before_leader_sync_total | counter | Log entries committed on followers' synced copies before the leader's own copy reached disk (leader_sends_before_sync in [cluster.consensus]). |
scramdb_cluster_replication_leader_unsynced_entries | gauge | Log entries the groups this node leads have sent but not yet synced to their own disk. |
scramdb_cluster_replication_lag_entries | gauge | Log entries a peer has not acknowledged, summed over the groups this node leads, labelled peer. A peer whose lag keeps growing is falling behind. |
scramdb_cluster_replication_leader_sync_seconds | histogram | Time from a leader writing new log entries to their sync to disk. |
scramdb_cluster_election_timeout_min_milliseconds | gauge | Lower bound of the election timeout in force (0 until a replication group has started on this node). |
scramdb_cluster_election_timeout_max_milliseconds | gauge | Upper bound of the election timeout in force. |
sends_paused_total climbing together with window_bytes{reason="ceiling"} toward a peer says the ceiling, not the link, is the limit: raise max_inflight_bytes for a long, fast link. Steady growth of transitions_total{cause="stall"} or {cause="reject"} on a healthy network says a follower keeps losing its place and is worth a look at its disk and its connection.
Consensus core buffers
Consensus reuses its working buffers across steps, so a steady stream of heartbeats, appends and commits allocates nothing for its step output; only a burst above the recent load, or the first load after a quiet spell, needs a fresh allocation, and that spare capacity is handed back again once the burst passes. A received frame is decoded without copying the commands and snapshot data it carries. These figures are for the whole process, summed across it.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_raft_step_vectors_reused_total | counter | Raft step output vectors served from a kept buffer instead of a new allocation. Counted as each vector is taken. |
scramdb_cluster_raft_step_vectors_fresh_total | counter | Raft step output vectors started without a kept buffer, so their first element allocates. Counted as each vector is taken. It stays flat on a steady load and grows when the load keeps more step output in flight than it recently did (a burst, or the first load after a quiet spell). |
scramdb_cluster_raft_step_buffer_bytes | gauge | Bytes the consensus threads keep for Raft step output. |
scramdb_cluster_raft_frame_blob_bytes_shared_total | counter | Command and snapshot bytes decoded from received Raft frames without a copy. |
scramdb_cluster_raft_frame_blob_bytes_copied_total | counter | Command and snapshot bytes copied out of received Raft frames. It stays at zero on a node: every receive path decodes from the frame it was given. |
scramdb_cluster_raft_frames_encoded_total | counter | Cluster messages encoded into a consensus thread's kept frame buffer. |
scramdb_cluster_raft_frame_buffer_growths_total | counter | Times a consensus thread's kept frame buffer grew to hold a message. |
scramdb_cluster_raft_frame_entry_vectors_decoded_total | counter | Entry vectors the Raft frame decoder allocated for received appends, one per append that carries entries. |
Heartbeats and quiet groups
A node sends one heartbeat frame per peer node each heartbeat interval, carrying a line for every group it leads with that peer, instead of one heartbeat per group ([cluster.consensus] coalesced_heartbeats). A frame names its group set by a hash and lists the groups in full only after the set changed. When a group's commit index advances, a read needs confirming or a small append waits, the frame leaves early: at once while the connection to that peer has nothing queued, otherwise after heartbeat_flush_delay. A small append rides inside the frame instead of taking a frame of its own (carry_small_appends), and a follower's reply to an append can ride back inside the answer frame it sends anyway. An idle group (nothing in flight, every member caught up) goes quiet: it stops its heartbeats and its members' election timers, and its peer is refreshed once per election_timeout_min (quiescence). See [cluster.consensus].
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_heartbeat_frames_sent_total | counter | Heartbeat frames this node sent, each covering every group it leads with that peer. |
scramdb_cluster_heartbeat_group_lines_sent_total | counter | Group lines the heartbeat frames of this node carried. |
scramdb_cluster_heartbeat_bytes_sent_total | counter | Bytes of heartbeat and answer frames this node sent. |
scramdb_cluster_heartbeat_full_lists_sent_total | counter | Heartbeat frames that listed every group this node leads with the peer, after that set changed. |
scramdb_cluster_heartbeat_quiet_frames_sent_total | counter | Heartbeat frames with no group line, sent to keep a peer's quiet groups quiet. |
scramdb_cluster_heartbeat_flushes_total | counter | Heartbeat frames sent early because a commit index advanced, a read needed confirming or a small append waited. |
scramdb_cluster_heartbeat_immediate_flushes_total | counter | Early heartbeat frames sent at once because the connection to the peer had nothing queued. |
scramdb_cluster_heartbeat_flush_delay_seconds | histogram | Time from a commit index advancing, or a read or small append waiting, to the heartbeat frame carrying it. |
scramdb_cluster_heartbeat_answers_sent_total | counter | Heartbeat answer frames this node sent to the nodes leading its groups. |
scramdb_cluster_heartbeat_frames_received_total | counter | Heartbeat frames this node received. |
scramdb_cluster_heartbeat_answers_received_total | counter | Heartbeat answer frames this node received. |
scramdb_cluster_heartbeat_round_trip_seconds | histogram | Time from sending a heartbeat frame to a peer to the peer's answer. |
scramdb_cluster_heartbeat_carried_appends_total | counter | Small appends sent inside heartbeat frames instead of frames of their own. |
scramdb_cluster_heartbeat_carried_bytes_total | counter | Bytes of the appends sent inside heartbeat frames. |
scramdb_cluster_heartbeat_append_replies_total | counter | Replies this node's groups sent to appends of the nodes leading them. |
scramdb_cluster_heartbeat_append_reply_frames_total | counter | Frames this node sent for its groups' append replies: a reply sent on its own, or an answer frame carrying replies and no heartbeat answer. |
scramdb_cluster_heartbeat_carried_replies_total | counter | Append replies sent inside heartbeat answer frames instead of frames of their own. |
scramdb_cluster_heartbeat_carried_replies_received_total | counter | Append replies this node received inside heartbeat answer frames. |
scramdb_cluster_heartbeat_send_refused_total | counter | Heartbeat and answer frames the transport refused because its queue was full or the peer was not connected. |
scramdb_cluster_heartbeat_delivery_refused_total | counter | Heartbeat lines a group could not take because its queue was full or it was gone. |
scramdb_cluster_heartbeat_malformed_frames_total | counter | Heartbeat frames dropped because they did not decode. |
scramdb_cluster_heartbeat_stale_views_total | counter | Heartbeat frames whose group-set hash did not match the set this node held, which asks for the full list. |
scramdb_cluster_heartbeat_quiesce_entered_total | counter | Groups this node leads that went quiet: every member caught up and idle, no heartbeats and no timers. |
scramdb_cluster_heartbeat_quiesce_left_total | counter | Quiet groups this node leads that woke up. |
scramdb_cluster_heartbeat_woken_by_set_change_total | counter | Quiet groups on this node woken because the node leading them stopped listing them. |
scramdb_cluster_heartbeat_woken_by_missed_refresh_total | counter | Quiet groups on this node woken because the node leading them missed two refreshes. |
scramdb_cluster_heartbeat_led_groups | gauge | Groups this node leads whose heartbeats its frames carry. |
scramdb_cluster_heartbeat_quiescent_groups | gauge | Groups this node leads that are quiet. |
scramdb_cluster_heartbeat_peers | gauge | Peer nodes this node sends heartbeat frames to or receives them from. |
scramdb_cluster_heartbeat_interval_milliseconds | gauge | The heartbeat interval in force (heartbeat_interval). |
scramdb_cluster_heartbeat_flush_delay_microseconds | gauge | How long an early heartbeat frame waits to gather more groups, in force (heartbeat_flush_delay). |
scramdb_cluster_heartbeat_quiet_refresh_milliseconds | gauge | How often a peer whose groups here are all quiet is refreshed, in force. |
group_lines_sent_total divided by frames_sent_total is how many groups one frame carries: the saving over one heartbeat per group. On an idle cluster quiescent_groups approaches led_groups and the heartbeat traffic falls to one small refresh per peer. woken_by_missed_refresh_total rising on a healthy network means a leader's node stalls; send_refused_total, delivery_refused_total and malformed_frames_total stay at zero on a healthy node.
Read-index forwarding
A replica that is not its group's leader, a follower or a learner, forwards its read-index requests to the leader, which answers them as it answers its own. A forwarded request with no answer within two election timeouts plus the round trip the replica measured to the leader, or whose connection to the leader closed, is refused. An answer that arrives after its request was refused still measures the round trip, so on a slow link the next request waits long enough; forward_late_total rising means a replica's link to its leader is slower than two election timeouts.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_read_index_forwarded_total | counter | Read-index requests this node forwarded to a group's leader. |
scramdb_cluster_read_index_forward_unanswered_total | counter | Forwarded read-index requests that ended without the leader's answer. |
scramdb_cluster_read_index_forward_late_total | counter | Leader answers to forwarded read-index requests that arrived after the request was given up. |
scramdb_cluster_read_index_served_for_peers_total | counter | Read-index requests this node answered as a group's leader for another replica. |
scramdb_cluster_read_index_forward_seconds | histogram | Time from forwarding a read-index request to the leader's answer. |
Requests to the metadata and shard groups
Every request a node sends its metadata group or one of its shard groups has a deadline adapted to the group's election timing. A request the group did not take in time changed nothing and is retried by its caller; a proposal whose answer did not come may still commit, so it is never sent again; a read confirmation is asked again until the deadline. A DDL statement forwarded to the metadata group's leader is bounded the same way. Process-wide, always present, all zero on a healthy cluster:
| Metric | Type | Meaning |
|---|---|---|
scramdb_group0_reply_timeouts_total | counter | Metadata group proposals whose answer did not come in time; each may still commit. |
scramdb_group0_mailbox_timeouts_total | counter | Metadata group requests the group did not take in time; none changed anything. |
scramdb_group0_read_retries_total | counter | Metadata group read confirmations asked again after an answer did not come. |
scramdb_group0_apply_wait_timeouts_total | counter | Waits for this node to apply a committed metadata group entry that ran out of time. |
scramdb_shard_group_reply_timeouts_total | counter | Shard group proposals whose answer did not come in time; each may still commit. |
scramdb_shard_group_mailbox_timeouts_total | counter | Shard group requests the group did not take in time; none changed anything. |
scramdb_shard_group_read_retries_total | counter | Shard group read points asked again after an answer did not come. |
scramdb_shard_group_apply_wait_timeouts_total | counter | Waits for this node to apply a committed shard group entry that ran out of time. |
scramdb_ddl_forward_answer_timeouts_total | counter | Forwarded metadata group proposals whose answer did not come in time. |
scramdb_ddl_forward_in_doubt_total | counter | Forwarded metadata group proposals whose leader could not learn their outcome. |
reply_timeouts_total or apply_wait_timeouts_total rising means a group is slow to commit or a node is slow to apply; mailbox_timeouts_total rising means the group itself is falling behind on incoming requests.
A metadata group proposal this node's answer did not reach in time is sent again under its first attempt's id, so a leader that already applied it once applies it once, never twice, and a resend past its own deadline is refused rather than applied late:
| Metric | Type | Meaning |
|---|---|---|
scramdb_group0_proposal_resends_total | counter | Metadata proposals sent again under the id of their first attempt. |
scramdb_group0_proposal_duplicates_total | counter | Metadata log entries applied as a repeat of a proposal already applied. |
scramdb_group0_proposal_expired_total | counter | Metadata proposals refused because they arrived after their deadline. |
scramdb_group0_proposal_window_ids | gauge | Metadata proposal ids this node keeps to answer a repeat. |
Read positions
Every read of a cluster table happens at a read position of the commit protocol. A fresh position first has each group's leader confirm the group's latest commit, brings a replica up to it, and waits out any commit still in flight on what is read, so a read never misses a commit already acknowledged to any client, on any node; a reader then serves each group it scans at that position before it reads. For a group this node holds no replica of, the position is asked of the group's leader first, then of its followers, and of its learners last: the leader has nothing to catch up, while a learner may trail its group by any amount, so a learner that has fallen behind does not hold up a statement while a voter of the group can answer. A replica that refuses the position (it does not run the group yet, it is starting, it could not catch up in time; counted on that node by scramdb_cluster_read_position_refusals_total), or fails while still connected, is passed over for the group's next replica. When no replica of such a group is connected to this node (a partition that just healed while its connections form again, a node restarting), the read waits for one to connect (scramdb_cluster_read_position_connect_waits_total). When every replica of the group has refused, as happens for a table created a moment before whose replicas are still starting, the read asks them all again after a pause of a few round trips that grows with each try, and serves the group itself once its own replica of it runs (scramdb_cluster_read_position_reask_waits_total). Either way the read waits for at most its read wait, and then fails with 40001 naming the group. Nothing here writes to a log unless a group is behind the position, where one fence entry brings it up (scramdb_cluster_dilith_fences_appended_total).
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_read_positions_total | counter | Fresh read positions chosen by this node. |
scramdb_cluster_read_position_seconds | histogram | Time to choose a fresh read position, from the first barrier to the last answer. |
scramdb_cluster_read_position_remote_groups_total | counter | Shard groups whose fresh read position another node took for this node (groups this node holds no replica of). |
scramdb_cluster_read_position_connect_waits_total | counter | Fresh read positions that waited for a replica of a shard group to connect. |
scramdb_cluster_read_position_reask_waits_total | counter | Fresh read positions that asked a shard group again after every replica refused it. |
scramdb_cluster_read_position_refusals_total | counter | Fresh read positions refused here because this node's replica could not serve them. |
scramdb_cluster_read_position_raised_total | counter | Transactions whose read position rose to reach a table they had not read before. |
scramdb_cluster_read_barrier_retries_total | counter | Read barriers asked again because a shard group had no leader to confirm them. |
scramdb_cluster_read_serves_total | counter | Shard groups served at a read position on this node. |
scramdb_cluster_read_serve_seconds | histogram | Time a shard group took to reach a read position on this node. |
scramdb_cluster_read_failures_total | counter | Reads refused because a shard group could not be confirmed or served in time. |
scramdb_cluster_read_registered_ranges_total | counter | Table reads registered with a transaction's commit for validation. |
scramdb_cluster_read_skip_serves_total | counter | Shard groups served for SKIP LOCKED or NOWAIT reads without waiting for a commit. |
scramdb_cluster_read_skip_waits_total | counter | SKIP LOCKED or NOWAIT reads of a shard group that waited for a commit in flight. |
scramdb_cluster_read_skip_fence_waits_total | counter | Waits of SKIP LOCKED or NOWAIT reads for a shard group to be fenced at their cut. |
scramdb_cluster_read_skipped_keys_total | counter | Keys a SKIP LOCKED or NOWAIT read found written by a commit in flight. |
scramdb_cluster_read_skipped_rows_total | counter | Rows SKIP LOCKED left out because a commit in flight writes them. |
scramdb_cluster_read_nowait_refusals_total | counter | NOWAIT statements refused because a commit in flight writes a row they select. |
scramdb_cluster_read_join_spills_total | counter | Reads over more shard groups than fit inline, whose answers spilled to the heap: reads that touch more than eight groups at once. |
scramdb_cluster_read_horizon_serves_total | counter | Shard group watermarks taken here for a transaction snapshot, waiting for no commit. |
scramdb_cluster_read_unreached_groups_total | counter | Shard groups a transaction snapshot could reach on no replica and left out. |
scramdb_cluster_read_fence_waits_total | counter | Waits for a shard group to be fenced at a read position, waiting for no commit. |
scramdb_cluster_read_range_serves_total | counter | Key ranges served at a read position that waited only for commits writing them. |
A read that finds no leader for a group keeps asking until its wait runs out (one election timeout plus the time a group may hold requests without a leader plus the time an undecided transaction may block it, about ten seconds with the defaults) and is then refused with SQLSTATE 40001; a replica that cannot catch up is refused with 57P03. Growth of read_failures_total together with read_barrier_retries_total means a group is without a leader. read_position_raised_total counts transactions that read a table only after an earlier statement had fixed their position: such a transaction's earlier reads are validated at their own position when it commits. Under READ COMMITTED it also counts statements whose position rose because a table they reached later, such as a foreign key's parent, was fresher, so a row committed before the statement began is never missed.
SKIP LOCKED and NOWAIT read the tables they lock without waiting for any commit in flight: the keys such commits write come back with the rows (read_skipped_keys_total), and the statement leaves those rows out (read_skipped_rows_total) or refuses with 55P03 (read_nowait_refusals_total). A shard group where a commit in flight writes a table the statement reads without locking it, such as a join's other side, is read the ordinary way, waiting for that commit, and counts in read_skip_waits_total.
A REPEATABLE READ or SERIALIZABLE transaction takes its snapshot from every shard group's confirmed watermark without waiting for any commit in flight (read_horizon_serves_total), then fences the groups this node holds at it (read_fence_waits_total counts the waits for a fence on its way). A read of a few keys at that snapshot waits only for a commit in flight on those keys (read_range_serves_total). A group the snapshot can reach on no replica, such as one cut off by a partition, is left out (read_unreached_groups_total): the transaction still reads every other table, and its reads are checked at COMMIT, which fails with 40001 if one of them missed a commit. A group whose replicas the snapshot reaches but which all refuse it, such as the groups of a table created a moment before whose replicas are still starting, is not left out: it is asked again at the horizon retry interval (horizon_retry, 90 ms by default) until a replica answers, for at most the read wait.
Adaptive routing
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_routes_total | counter/label | Routing decisions, labelled route: one per statement, and one per scatter for a statement that scatters more than once (each outer row of a k-NN LATERAL join). |
route takes exactly one of five values: oltp, point, local_replica, forward_replica, scatter. That's the whole label space, fixed by construction, the keys are the five route kinds the planner can choose, never a node id or a group id, so this metric's cardinality can never grow with the size of the cluster. A route kind the counter doesn't recognize is dropped rather than bucketed into a sixth catch-all series, so a routing path that forgot to name itself shows up as a silent gap here, not a mislabeled entry.
Where fragments run
A distributed read runs each bucket's fragment on one replica that holds the bucket: a learner before a follower before the shard group's leader, in this node's region first, spread evenly (see Where fragments run). [cluster] fragment_any_replica = false keeps fragments on the buckets' owners.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_exchange_placed_learner_buckets_total | counter | Buckets of distributed reads this node placed on a learner replica. |
scramdb_cluster_exchange_placed_follower_buckets_total | counter | Buckets of distributed reads this node placed on a follower replica. |
scramdb_cluster_exchange_placed_leader_buckets_total | counter | Buckets of distributed reads this node placed on a shard group's leader. |
scramdb_cluster_exchange_placement_skips_total | counter | Replicas placement passed over because they could not serve the read position. |
scramdb_cluster_exchange_fragment_redispatches_total | counter | Fragments run on another replica after one could not serve them. |
scramdb_cluster_exchange_fragments_served_learner_total | counter | Fragments this node served from its learner store. |
scramdb_cluster_exchange_round_replacements_total | counter | Shuffle and two-phase read rounds run again on other replicas after a node could not serve its buckets. |
scramdb_cluster_exchange_fragments_refused_total | counter | Fragments this node refused because its replica could not serve the read position. |
scramdb_cluster_exchange_shipped_tables_total | counter | Tables this node sent with a query's fragments because no replica held them whole. |
placed_leader_buckets_total growing against the learner and follower counts means the learners and followers cannot serve the reads (behind their group, or waiting for a full copy of its rows), so analytical scans land on the leaders; fragments_refused_total that stays high on one node says that node's replicas lag.
Distributed shuffle joins
How a shuffle join across nodes adapted while it ran ([cluster.distributed_join] in the config file).
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_shuffle_consumer_fragments_total | counter | Consumer fragments dispatched for a materialized shuffle join, one per coalesce group. Equals the participant count with coalescing off; lower with it on. Zero on a streaming query. |
scramdb_cluster_shuffle_aqe_coalesce_groups_total | counter | Coalesce groups actually acted on: each merged more than one small adjacent partition into a single consumer fragment. |
scramdb_cluster_shuffle_aqe_skew_split_flagged_total | counter | Partitions flagged skewed (over 5x the median and over 256MB) but not acted on: skew-split off, or a single-node participant set with nothing to split across. Logged and metered even when nothing follows. |
scramdb_cluster_shuffle_aqe_skew_split_acted_total | counter | Sub-consumers dispatched for a skew-split partition that was acted on: the larger side split into disjoint row ranges, the smaller side replicated. Summed across every acted partition. |
scramdb_cluster_shuffle_aqe_broadcast_demote_flagged_total | counter | Broadcast sides flagged over budget, a broadcast-demote candidate. Logged and metered, not acted on yet. |
scramdb_cluster_shuffle_straggler_backups_total | counter | Speculative backup consumer fragments dispatched because a primary consumer didn't return within the straggler-backup delay, on a different live node reading the same frozen spools. |
Node operations and cluster TRUNCATE
The whole-cluster vector statistics views and a cluster TRUNCATE reach every other node
through the same exchange the distributed queries use.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_exchange_node_ops_sent_total / scramdb_cluster_exchange_node_ops_failed_total | counter | Node operations this node asked a peer to run, and those a peer could not run or never answered. |
scramdb_cluster_exchange_node_ops_served_total / scramdb_cluster_exchange_node_ops_serve_failed_total | counter | Node operations this node ran for a peer, and those it could not. |
scramdb_cluster_exchange_view_gathers_total / scramdb_cluster_exchange_view_nodes_unreachable_total | counter | Whole-cluster statistics view reads this node gathered, and nodes such a read reported as unreachable. |
scramdb_cluster_exchange_view_gather_seconds | histogram | Time to gather one whole-cluster statistics view read. |
scramdb_cluster_exchange_truncates_total / scramdb_cluster_exchange_truncate_failures_total | counter | Cluster TRUNCATE statements this node coordinated, and those that did not complete on every node. |
scramdb_cluster_exchange_truncate_seconds | histogram | Time for one cluster TRUNCATE to reach every node holding the table. |
scramdb_cluster_exchange_producer_failures_propagated_total / scramdb_cluster_exchange_producer_failure_undelivered_total | counter | Failed shuffle producers that handed their error to every consumer at once, and consumers they could not reach (those fall back to their own receive deadline). |
Cluster branches and table commands
A CREATE DATABASE ... CLONE or a scram.branch_at on a cluster, and the table commands that
carry a branch, a CLONE's fork hold and a cluster TRUNCATE through each shard group's log.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_branches_created_total | counter | Branch databases this node created on the cluster. |
scramdb_cluster_branch_failures_total | counter | Branch databases this node could not create on the cluster. |
scramdb_cluster_branch_tables_forked_total | counter | Tables forked into a branch database on this node's stores. |
scramdb_cluster_branch_pending_rows_total | counter | Rows a shard group gave a branch after this node's store forked it, before the group's fork entry. |
scramdb_cluster_branch_shared_stages_total | counter | COPY segments shared into a branch as they attached after this node's store forked it. |
scramdb_cluster_branch_holds_taken_total | counter | Fork holds taken on this node's stores, each keeping row versions until its branch is forked. |
scramdb_cluster_branch_holds_released_total | counter | Fork holds released on this node's stores. |
scramdb_cluster_branch_clones_abandoned_total | counter | CLONE statements this node abandoned, their fork holds released on every node. |
scramdb_cluster_branch_archive_builds_total | counter | Branches at a past time this node built from the cluster archive. |
scramdb_cluster_branch_archive_rows_total | counter | Rows this node read from the cluster archive into a branch at a past time. |
scramdb_cluster_dropped_database_rows_skipped_total | counter | Rows of a dropped database a shard group skipped on this node. |
scramdb_cluster_branch_fork_records | gauge | Branch fork entries the shard groups on this node still record one by one. A record is folded away once the store holding its group settled the branch (made it ready, or dropped it), so this stays at the branches in progress. |
scramdb_cluster_branch_fork_records_pruned_total | counter | Branch fork records folded into a group's floor once their branch was settled. |
scramdb_cluster_dropped_databases_recorded | gauge | Dropped databases this node's stores record one by one; each is released once its drop is decided. |
scramdb_cluster_table_commands_proposed_total | counter | Table commands this node placed in a shard group's log. |
scramdb_cluster_table_command_failures_total | counter | Table commands this node could not place in a shard group's log before the deadline. |
scramdb_cluster_table_command_redirects_total | counter | Table command proposals a node turned away because its replica did not lead the group. |
Table definition changes
An ALTER TABLE on a cluster: the commits it holds back on each node, the fence it places in each of a distributed table's shard groups, and the rows each node installs across a change.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_schema_changes_pending | gauge | Tables whose definition change is under way on this node's stores. |
scramdb_cluster_schema_commits_refused_total | counter | Commits refused because a table they wrote changed its definition while they ran. |
scramdb_cluster_schema_commit_drains_total | counter | Waits for the commits under way before a table's definition change. |
scramdb_cluster_schema_commit_drain_timeouts_total | counter | Waits for the commits under way that ended at their deadline. |
scramdb_cluster_schema_fences_applied_total | counter | Table definition change fences this node's stores applied from shard group logs. |
scramdb_cluster_schema_fence_waits_total | counter | Fences that waited for this node's catalog to take the change first. |
scramdb_cluster_renamed_table_rows_total | counter | Rows of a renamed distributed table installed under its new name. |
scramdb_cluster_widened_rows_total | counter | Rows written before a column was added, widened to the table's shape as installed. |
A DROP TABLE of a distributed table holds back and drains the table's commits and fences its shard groups the same way before the table goes, so it moves schema_commit_drains_total and schema_fences_applied_total too.
UPDATE and DELETE on a partially held table
An UPDATE or DELETE on a node that holds only some buckets of a distributed table reads its
rows from every owner (see
UPDATE and DELETE across nodes).
A node that holds every bucket of the table runs the single-node path and moves none of these.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_dml_statements_distributed_total | counter | UPDATE and DELETE statements that read their rows from every owner. |
scramdb_cluster_dml_candidates_remote_total / scramdb_cluster_dml_candidates_local_total | counter | Rows such statements received from other nodes, and read from this node's own buckets. |
scramdb_cluster_dml_candidate_batches_total | counter | Row batches such statements processed. |
scramdb_cluster_dml_candidate_fetch_seconds | histogram | Time a statement waited for its next batch of rows. |
scramdb_cluster_dml_candidate_spilled_bytes_total | counter | Row bytes set aside on disk because the statement's memory budget was spoken for. |
scramdb_cluster_dml_owners_pruned_total | counter | Owners a statement skipped because its WHERE names one bucket. |
scramdb_cluster_dml_locking_selects_total | counter | Locking SELECT statements (FOR UPDATE, FOR NO KEY UPDATE, FOR SHARE, FOR KEY SHARE) that read a distributed table's rows without locking them. |
scramdb_cluster_dml_locking_read_rows_total | counter | Rows such statements took and registered one by one for their commit to check (a SERIALIZABLE statement without SKIP LOCKED has its whole scan checked instead and adds nothing; rows SKIP LOCKED left out are not counted). |
scramdb_cluster_dml_own_writes_reads_total | counter | Queries that read a partially held table together with their transaction's own uncommitted changes to it. |
scramdb_cluster_dml_candidate_readers_replaced_total | counter | Readers of an UPDATE, DELETE or uniqueness check that could not serve their buckets and were read again on another replica. |
A growing candidate_spilled_bytes_total means large UPDATE or DELETE statements are running
close to the execution memory budget.
The row-lock family these statements used to report (scramdb_cluster_dml_row_lock_batches_total,
scramdb_cluster_dml_row_locks_total, scramdb_cluster_dml_row_lock_waits_total,
scramdb_cluster_dml_row_lock_wait_seconds, scramdb_cluster_dml_row_lock_conflicts_total,
scramdb_cluster_dml_row_locks_released_total) and scramdb_cluster_dml_statement_restarts_total
are gone: rows of a distributed table are no longer locked, so no statement waits, restarts or
releases a lock (see
Row locks on distributed tables).
Writers that reach the same rows now show up as 40001 failures at commit: watch
scramdb_cluster_dilith_aborts_total and your clients' retry counts instead, and remove any alert
or dashboard panel that reads the removed names.
Distributed statements
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_read_stability_retries_total | counter | Reads re-run because a replica applied writes while they scanned. |
scramdb_cluster_read_stability_failures_total | counter | Reads that failed because replicas kept applying writes through every re-run. The statement fails with 40001 and can be retried. |
scramdb_cluster_serialization_retries_total | counter | Autocommit statements sent again, and COPY votes voted again, after a serialization failure. |
scramdb_cluster_placement_waits_total | counter | Statements that waited for this node to learn where a table they name is placed, right after it restarted or right after the table was created. |
scramdb_cluster_rows_moved_total | counter | Rows an UPDATE moved to a new primary key or shard group. |
scramdb_cluster_unique_probes_total | counter | Uniqueness checks of new keys sent to the nodes that hold their buckets. |
scramdb_cluster_unique_violations_total | counter | New keys refused because a row on another node already holds them. |
scramdb_cluster_inserts_positioned_by_key_total | counter | Inserted batches whose keys were checked in their own shard groups only, not in every group of the table. |
scramdb_cluster_upsert_remote_conflicts_total | counter | INSERT ... ON CONFLICT rows whose conflicting row is stored on another node. |
scramdb_cluster_on_conflict_keys_overtaken_total | counter | INSERT ... ON CONFLICT statements refused because another transaction committed one of their conflict keys after the statement judged it free. |
scramdb_cluster_own_writes_spilled_bytes_total | counter | Bytes of a transaction's view of a table a query set aside on disk for want of memory. |
scramdb_cluster_general_statements_total | counter | Statements answered on their coordinator over rows streamed from every shard group. |
scramdb_cluster_general_fast_path_declines_total | counter | Statements a faster distributed path handed to the general path. |
scramdb_cluster_general_rows_streamed_total | counter | Rows streamed to a coordinator for the statements it answered over them. |
scramdb_cluster_general_spilled_bytes_total | counter | Bytes of streamed rows a coordinator set aside on disk for want of memory. |
scramdb_cluster_general_readers_refused_total | counter | Readers of a streamed relation whose replica refused the read position. Their buckets were read on another replica. |
scramdb_cluster_general_readers_unreached_total | counter | Readers of a streamed relation that were not reached or stopped sending. Their buckets were read on another replica. |
scramdb_cluster_general_readers_failed_total | counter | Readers of a streamed relation whose read failed, failing the statement. |
scramdb_cluster_general_reads_exhausted_total | counter | Streamed reads that failed because no replica was left to read a bucket. The statement fails with 40001 and can be retried. |
Prepared transactions
PREPARE TRANSACTION, COMMIT PREPARED and ROLLBACK PREPARED on a cluster's distributed tables (see Two-phase commit), counted on the node that prepared the transaction.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_two_phase_prepared_total | counter | Transactions prepared on this node. |
scramdb_cluster_two_phase_prepare_failures_total | counter | PREPARE TRANSACTION statements that failed and rolled their transaction back. |
scramdb_cluster_two_phase_committed_total | counter | Prepared transactions committed on this node. |
scramdb_cluster_two_phase_commit_conflicts_total | counter | COMMIT PREPARED statements refused with 40001 because a later commit overtook the transaction. |
scramdb_cluster_two_phase_commit_unknown_total | counter | COMMIT PREPARED statements whose outcome was not known when they answered (08007). |
scramdb_cluster_two_phase_commit_failures_total | counter | COMMIT PREPARED statements that failed for another reason, such as an unknown transaction name. |
scramdb_cluster_two_phase_placement_waits_total | counter | COMMIT PREPARED statements that waited, right after this node restarted, for it to learn where their tables are placed. |
scramdb_cluster_two_phase_rolled_back_total | counter | Prepared transactions rolled back on this node. |
scramdb_cluster_dilith_prepared_attempts_total | counter | Commit attempts PREPARE TRANSACTION kept for a later COMMIT PREPARED. |
scramdb_cluster_dilith_prepared_attempts_resumed_total | counter | Kept attempts COMMIT PREPARED took up to commit. |
A growing commit_conflicts_total means prepared transactions wait long enough for other writers to change what they read; commit them sooner, or retry them from the start.
Columnar mirror
A node that keeps a columnar copy of the rows it replicates (a hybrid voter, columnar_replica = true, or a learner) copies committed entries into it on its own mirror threads.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_columnar_mirror_threads | gauge | Threads that copy committed entries into this node's columnar store. |
scramdb_cluster_columnar_mirror_queued | gauge | Columnar mirror lanes waiting for a mirror thread. |
scramdb_cluster_columnar_mirror_panics_total | counter | Columnar mirror lanes stopped by a panic. Zero on a healthy node; the log names the cause. |
Learner-read freshness
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_learner_freshness_seconds | histogram | How long ago each queried group's read floor last advanced, observed once per stale or strong learner_read read. The floor is the commit protocol's closed position at the replica: every commit at or below it is installed there and none can land there later. The OLTP path (learner_read off) never touches this histogram. Same buckets as the handshake histogram, 1ms to 5s plus +Inf, a natural fit for replication lag. |
This is the same lag SHOW FRESHNESS reports per data group in Clustering, captured here as a scrapeable distribution instead of a point-in-time value.
Faults, retries and reachability
How a node sees its peers and its consensus groups, and every retry, timeout and re-dispatch of the cluster's network mechanisms. Always emitted in cluster mode. On a healthy cluster the failure and retry counters stay at zero, and the gauges show the peers and members that are up. A fault that repeats is logged once, as a warning when it begins, and counted every time; a lost peer that is heard again is logged once at info.
| Metric | Type | Meaning |
|---|---|---|
scramdb_cluster_peers_reachable | gauge | Peers with an open cluster connection. |
scramdb_cluster_peer_reachable | gauge | Whether a peer connected or lost now has an open cluster connection (1 or 0), labelled peer. |
scramdb_cluster_peer_round_trip_microseconds | gauge | Round trip of the newest answered liveness ping to a peer, labelled peer. Absent until a ping is answered. |
scramdb_cluster_peer_silence_milliseconds | gauge | Time since a peer's cluster connection last received a byte, labelled peer. |
scramdb_cluster_peers_lost | gauge | Peers that were connected and have not been heard from since their connection was lost. |
scramdb_cluster_peer_losses_total | counter | Times a connected peer became unreachable. A partition is one loss per peer, however often its connections close. |
scramdb_cluster_dial_failures_total | counter | Connection attempts to peers that failed, labelled cause. |
scramdb_cluster_inbound_handshake_failures_total | counter | Connections from peers that failed or were refused during the handshake, labelled cause. |
scramdb_cluster_accept_errors_total | counter | Times a cluster listener failed to accept a connection. |
scramdb_cluster_groups_without_leader | gauge | Consensus groups on this node that know no leader. |
scramdb_cluster_group_leader_losses_total | counter | Times a consensus group on this node lost its leader. |
scramdb_cluster_stalled_groups | gauge | Consensus groups on this node that made no progress for as long as a request waits for its answer (fifteen election timeouts, at least 5 s). |
scramdb_cluster_group_stalls_total | counter | Times a consensus group on this node stalled. |
scramdb_cluster_retries_total | counter | Times a cluster mechanism tried again after a failure, labelled mechanism. |
scramdb_cluster_retries_exhausted_total | counter | Times a cluster mechanism gave up after its retries or its deadline ran out, labelled mechanism. |
scramdb_cluster_timeouts_total | counter | Waits of a cluster mechanism that reached their deadline, labelled mechanism. |
scramdb_cluster_redispatches_total | counter | Work placed again on another node after its first node failed, labelled mechanism. |
scramdb_cluster_retry_rate_per_second | gauge | Retries a second across every mechanism over the last ten seconds. A sustained rise is a retry storm. |
scramdb_group0_reply_timeouts_total / scramdb_shard_group_reply_timeouts_total | counter | Metadata group and shard group proposals whose answer did not come in time. Each may still commit. |
scramdb_group0_mailbox_timeouts_total / scramdb_shard_group_mailbox_timeouts_total | counter | Requests the group did not take in time. None changed anything. |
scramdb_group0_read_retries_total / scramdb_shard_group_read_retries_total | counter | Read confirmations asked again after an answer did not come. |
scramdb_group0_apply_wait_timeouts_total / scramdb_shard_group_apply_wait_timeouts_total | counter | Waits for this node to apply a committed entry that ran out of time. |
scramdb_cluster_swim_members_alive / _suspect / _dead | gauge | Cluster members this node's failure detector sees alive, suspects, and has declared failed. |
scramdb_cluster_swim_local_health | gauge | This node's failure detector health score. Above zero, the node itself is missing probe answers. |
scramdb_cluster_swim_member_up_total / _suspect_total / _down_total | counter | Times a member came up, became suspect, or was declared failed in this node's failure detector. |
scramdb_cluster_swim_malformed_messages_total | counter | Failure detector messages from peers that could not be read and were dropped. |
scramdb_cluster_swim_send_failures_total | counter | Failure detector messages this node could not hand to a peer's connection. |
scramdb_cluster_swim_joins_dropped_total | counter | Member joins dropped because the failure detector's join queue was full. |
scramdb_cluster_exchange_partial_chunks | gauge | Exchange streams holding part of a chunk in reassembly on this node. |
scramdb_cluster_exchange_malformed_frames_total / scramdb_cluster_exchange_orphan_frames_total | counter | Exchange frames dropped because they could not be read, or because no query on this node was waiting for their stream. |
scramdb_cluster_exchange_out_of_order_total / _undecodable_chunks_total / _reassembly_overflows_total | counter | Exchange streams failed because a frame arrived out of order, a chunk could not be decoded, or a chunk grew past the reassembly limit. |
scramdb_cluster_exchange_producer_failures_received_total | counter | Exchange streams whose producer reported a failure to this node. |
scramdb_cluster_stream_frames_malformed_total | counter | Frames of a cluster service from peers that could not be read and were dropped. |
scramdb_cluster_point_forwards_total | counter | Point reads this node forwarded to their group's leader. |
scramdb_cluster_point_forward_send_failures_total / _remote_errors_total / _reply_send_failures_total | counter | Forwarded point reads that could not be sent, that the leader answered with an error, and answers to other nodes' forwarded reads that could not be sent back. A forwarded read that could not be sent or got no answer is answered on the general path instead; an error the leader reported is the statement's answer. |
The cause of a failed connection is one of connect_timeout, refused, reset, unreachable, unresolved, handshake_timeout, rejected, version (the two nodes share no cluster protocol version: upgrade the older one), identity, tls, protocol, closed and other.
The mechanism is one of the following, and each family lists only the mechanisms that can record it, so no sample reads as a measured zero:
dialandlane_dial: connecting to a peer;point_forward,fragmentandfetch: a distributed statement's requests;install: fetching a state snapshot;drainandautojoin: membership changes;metadata_groupandshard_group: requests to a consensus group;commit_voteandcommit_outcome: the commit protocol's resends.
A fragment timeout is a reader that missed its deadline. Its buckets are read on another replica, which is a fragment re-dispatch.
Alerting on cluster metrics
Three conditions worth paging on. Each maps to a documented design decision rather than a guessed threshold:
groups:
- name: scramdb-cluster
rules:
- alert: ScramDBGroupsQuarantined
expr: scramdb_cluster_quarantined_groups > 0
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $value }} raft group(s) quarantined"
description: >
A group's storage was quarantined at boot for corruption and has
not finished its membership-replace rejoin. This gauge is
recomputed from the real gated set every reconcile round, not
decremented, so restarting the node will not clear a stale
reading; only a completed rejoin does.
- alert: ScramDBLogCompactionStalled
expr: |
rate(scramdb_cluster_log_compactions_total[15m]) == 0
and rate(scramdb_cluster_group_commit_ops_total[15m]) > 0
for: 30m
labels:
severity: warning
annotations:
summary: "Raft log compaction has not run on a busy node in 30 minutes"
description: >
The node is committing but not compacting. Either its state
snapshots keep failing (check
scramdb_cluster_state_snapshot_failures_total and
scramdb_cluster_state_snapshots_abandoned_total: a full disk or a
memory_bytes budget too small), or a replica is behind but still
inside the lag limit. A replica lagging past the limit stops
holding the log and is repaired from a snapshot when it returns.
- alert: ScramDBPartitionDropsDetected
expr: increase(scramdb_cluster_partition_drops_total[5m]) > 0
labels:
severity: critical
annotations:
summary: "Cluster transport frames are being silently dropped"
description: >
This counter only moves when a network partition is injected
(chaos testing) or something outside the engine is dropping
cluster transport frames. It is zero in every normal deployment,
so any increase is worth paging on.
GIN index metrics
Process-wide totals over every GIN index (USING gin) on the node, always present, on the
same /metrics body. Per-index size is not labelled here; EXPLAIN shows which index a
query probes.
| Metric | Type | Meaning |
|---|---|---|
scramdb_gin_probes_total | counter | GIN index probes run. |
scramdb_gin_candidates_total | counter | Candidate rows the posting lists offered. |
scramdb_gin_recheck_rejected_total | counter | Candidates the recheck of the query's filter rejected. |
scramdb_gin_rows_returned_total | counter | Rows GIN probes returned after their recheck. |
scramdb_gin_snapshot_scans_total | counter | GIN-planned reads inside a REPEATABLE READ or SERIALIZABLE transaction, answered by scanning the table as of the snapshot. |
scramdb_gin_entries_inserted_total / scramdb_gin_entries_deleted_total | counter | Index entries added / removed by writes. |
scramdb_gin_probe_seconds | histogram | Time one GIN probe took, posting lists to returned rows. |
A high recheck_rejected_total against candidates_total means the index narrows poorly
for your filters; a rising snapshot_scans_total means snapshot transactions are reading
whole tables where a READ COMMITTED read would probe the index.
Index builds and compaction
An index build and a compaction merge never work on one table at the same time, as in
PostgreSQL, where CREATE INDEX and VACUUM take conflicting locks. Background compaction
yields: a merge running when a build starts stops, and the table is merged on a later pass. An
explicit VACUUM waits for the builds on its table to finish, and a build waits for a running
VACUUM, for at most lock_timeout when it is set (then it fails with SQLSTATE 55P03).
Process-wide, always present:
| Metric | Type | Meaning |
|---|---|---|
scramdb_compaction_merges_yielded_to_index_builds_total | counter | Background merges stopped or skipped because an index build held their table. |
scramdb_index_builds_waited_for_merges_total | counter | Index builds that waited for a merge running on their table to end. |
scramdb_index_build_merge_wait_timeouts_total | counter | Index builds that stopped waiting for a merge on their table at their lock timeout. |
scramdb_vacuums_waited_for_index_builds_total | counter | VACUUM runs that waited for the index builds holding their table to end. |
A separate fence orders a compaction's row relocation against a concurrent write that resolves a
row's physical position by table (DELETE, UPDATE, ON CONFLICT DO UPDATE/DO SELECT, a
replica's apply): the write holds the fence until its commit is durable, and a compaction raises
it before it swaps its merged segments in, so a swap never relocates a row out from under a write
already under way. A statement that reads a table's row identities (a ctid column, a
TABLESAMPLE draw) holds the same fence until it ends, or inside a transaction block until the
block ends, so each row keeps one ctid for it, as PostgreSQL's VACUUM FULL waits for every
transaction that read the table. Process-wide, always present:
| Metric | Type | Meaning |
|---|---|---|
scramdb_writes_waited_for_compaction_total | counter | Statements that waited for a compaction to finish moving their table's rows. |
scramdb_write_compaction_wait_timeouts_total | counter | Statements that stopped waiting for a compaction at their lock timeout. |
scramdb_compaction_guarded_swaps_total | counter | Compaction swaps made while the statements using their table's rows were held back. |
scramdb_compaction_swaps_abandoned_total | counter | Background compaction batches dropped because the statements holding their table's rows did not finish in time. |
scramdb_vacuum_lock_timeouts_total | counter | VACUUM FULL runs stopped at their lock timeout while statements held the table's rows. |
scramdb_compaction_late_deletes_carried_total | counter | Rows deleted while a compaction ran that it carried to their new places. |
scramdb_compaction_writer_drain_seconds | histogram | Time a compaction waited for the statements holding its table's rows to finish. |
scramdb_compaction_writer_pause_seconds | histogram | Time a compaction held back the statements using its table's rows while it swapped. |
scramdb_write_compaction_wait_seconds | histogram | Time a statement waited for a compaction before it could read or write its rows. |
scramdb_compaction_writer_tables | gauge | Tables with writer counters allocated. |
swaps_abandoned_total or vacuum_lock_timeouts_total rising means the statements holding a
table's rows (its writers, and its ctid or TABLESAMPLE readers) never drain within the wait
bound: background compaction backs off quietly, VACUUM FULL fails loudly with 55P03.
Query worker pool
The engine's core-pinned query worker pool ([execution] workers), process-wide, always present:
| Metric | Type | Meaning |
|---|---|---|
scramdb_query_worker_count | gauge | Pinned workers the engine's query worker pool runs. |
scramdb_query_worker_queue_depth | gauge | Outstanding tasks queued across the query worker pool's pinned workers, summed. |
scramdb_query_worker_cpu_quota_cores_permille | gauge | This engine's claimed cores per thousand of the machine's full, unconstrained cpuset: 1000 unconstrained, lower under a CPU limit (a container --cpus, a Kubernetes CPU request). |
worker_queue_depth climbing while worker_count holds steady means query work is arriving
faster than the pool drains it; cpu_quota_cores_permille well below 1000 on a box that looks
otherwise idle means the CFS quota, not the workload, is the ceiling on worker_count.
Stand-in query workers
A fixed-pool query worker whose core is parked on a lock wait hands its core to a stand-in worker for the wait's duration, so a lock contended by one statement does not idle every query on that core. Process-wide, always present:
| Metric | Type | Meaning |
|---|---|---|
scramdb_query_standins | gauge | Stand-in query worker threads alive right now, running or idle-cached. |
scramdb_query_standin_starts_total | counter | Stand-in query worker threads started. |
Vector index metrics
Every anode/manode vector index family lives on the scramdb_vector_ prefix,
appended to the same /metrics body as everything above. One sample feeds both the
Prometheus families below and the system views in the next section, so a number never
disagrees between them. A
family the sample has not measured (a recall job that never ran, a per-index rate with
no traffic yet) is omitted from the scrape entirely rather than rendered as a
fabricated 0, the same rule the WAL archive metrics above follow. Cardinality is
bounded by the number of indexes (and cluster parts) open on the node, reported by
scramdb_vector_series so you can see the cost.
Node totals (no index label, always present):
| Metric | Type | Meaning |
|---|---|---|
scramdb_vector_memory_budget_total_bytes | gauge | The memory ceiling configured for all vector indexes on this node together. |
scramdb_vector_memory_reserved_bytes | gauge | Memory currently reserved by vector index arenas against the total ceiling. |
scramdb_vector_memory_in_use_bytes | gauge | Memory vector indexes use right now: their arenas, row maps and write buffers together. The write buffers are charged to execution memory, not to the vector ceiling. A build's working memory (the model training sample, the convert buffers) is not counted here: it is charged to storage.memory.maintenance_bytes, like every index build. |
scramdb_vector_memory_engine_bytes | gauge | Memory the engine holds for vector index ledgers, bitmaps and buffers. |
scramdb_vector_memory_budget_refusals_total | counter | Index builds, rebuilds and arena regrows refused because the total ceiling would be exceeded. |
scramdb_vector_indexes | gauge | Vector indexes open on this node. |
scramdb_vector_regrows_total / scramdb_vector_regrows_running | counter / gauge | Arena regrows completed / running on this node. |
scramdb_vector_scans_total{tier} | counter | Indexed vector searches served, labelled exact, traversal, or postfilter. |
scramdb_vector_scan_seconds | histogram | Time an indexed vector search took end to end. |
scramdb_vector_rerank_rows_total / scramdb_vector_rerank_seconds | counter / histogram | Candidate rows re-ranked exactly after an index search, and how long it took. |
scramdb_vector_widening_rounds_total | counter | Extra index rounds taken when a filtered search had fewer than the requested rows. |
scramdb_vector_underfilled_total | counter | Filtered searches that returned fewer rows than requested after the scan limit. |
scramdb_vector_knn_join_batches_total / scramdb_vector_knn_join_queries_total | counter | Batches / individual query vectors sent to an index by a nearest-neighbour join. |
scramdb_vector_threshold_scans_total | counter | Distance threshold searches served by an index. |
scramdb_vector_exact_scans_total / scramdb_vector_exact_scan_seconds | counter / histogram | Nearest-neighbour statements served by the exact scan without an index, and how long it took. |
scramdb_vector_distributed_scans_total / scramdb_vector_distributed_merge_rows_total | counter | Nearest-neighbour statements fanned out to cluster readers, and rows merged from them. |
scramdb_vector_distributed_redispatches_total | counter | Readers of a distributed nearest-neighbour search whose buckets another replica served, after the reader missed its deadline or could not be reached. |
scramdb_vector_distributed_failures_total | counter | Distributed nearest-neighbour searches that failed because a reader failed or no replica of a bucket was left. |
scramdb_vector_builds_running / scramdb_vector_build_seconds | gauge / histogram | Vector index builds and rebuilds running, and how long one takes. |
scramdb_vector_reconcile_rows_total | counter | Rows replayed into vector indexes at startup to catch up with the table. |
scramdb_vector_build_owed_deletes_total | counter | Deleted rows an index build kept for read snapshots older than the delete, so those snapshots still find them. |
scramdb_vector_delete_tail_reclaims_total | counter | Reclaims that rebuilt a vector index's delete tail (the deleted rows kept for older snapshots) once some of those rows were no longer needed. |
scramdb_vector_delete_tail_batches_dropped_total / scramdb_vector_delete_tail_batches_relinked_total | counter | Delete tail batches a reclaim dropped whole because no read snapshot needs them, and batches it kept whole without copying them. |
scramdb_vector_delete_tail_batches_merged_total / scramdb_vector_delete_tail_boundary_batches_folded_total | counter | Small delete tail batches a reclaim copied into a larger one, and batches it read row by row because a read snapshot's boundary cut through them. Every batch a reclaim judges lands in exactly one of these four counters. |
scramdb_vector_fit_sample_bytes_total | counter | Bytes of training samples vector index model fits charged to execution memory: a build's fits through storage.memory.maintenance_bytes, every other fit (a streaming INSERT or COPY that trains the model, a background retrain) to the execution pool directly. |
scramdb_vector_setting{key} / scramdb_vector_setting_source{key,source} | gauge | The value in force for a [vector] setting, and which layer set it (config, override, setpoint). |
scramdb_vector_series | gauge | Prometheus series the vector index families currently emit, computed from the same scrape it appears in. |
Per index (labels index, table, column, method, group):
| Metric | Type | Meaning |
|---|---|---|
scramdb_vector_index_state{state} | gauge | Whether the index is in this state (building, ready, regrowing, rebuilding, tail, invalid); one state reads 1. |
scramdb_vector_index_live_vectors / scramdb_vector_index_tombstones | gauge | Vectors the index holds and can return, and deleted vectors it still carries until compaction. |
scramdb_vector_index_unindexed_rows | gauge | Rows served exactly because the index has no memory left to hold them (the tail state). |
scramdb_vector_index_ledger_live_rows / scramdb_vector_index_delete_tail_rows | gauge | Rows the index's row map holds that are not deleted (what a rebuild sizes the index for), and deleted rows the index keeps for read snapshots older than the delete until a reclaim finds no snapshot needs them. |
scramdb_vector_index_arena_slots / scramdb_vector_index_arena_used_permille | gauge | Slots the index's arena is promised, and the share in use, in thousandths. |
scramdb_vector_index_regrows_total / scramdb_vector_index_regrow_running | counter / gauge | Times the arena was reopened with a larger promise, and whether it is happening now. |
scramdb_vector_index_memory_budget_bytes / scramdb_vector_index_memory_in_use_bytes | gauge | The memory ceiling the index runs under, and resident memory it uses right now. |
scramdb_vector_index_memory_plane_bytes{plane} | gauge | Resident memory by plane: codes, graph, exact_vectors, directory, slot_state, coarse_model, journal_buffers. |
scramdb_vector_index_ledger_bytes / scramdb_vector_index_disk_bytes / scramdb_vector_index_journal_bytes | gauge | Memory the row-id ledger uses, bytes the index occupies on disk, and bytes of write history kept before compaction. |
scramdb_vector_index_ops_total{op} / scramdb_vector_index_ops_failed_total{op} | counter | Operations completed / refused or failed, op one of search, insert, update, delete, flush. |
scramdb_vector_index_search_qps | gauge | Searches per second over the recent window. |
scramdb_vector_index_search_latency_seconds{quantile} / scramdb_vector_index_write_latency_seconds{quantile} | gauge | Search / write latency over the recent window, quantile one of 0.5, 0.95, 0.99, 0.999, max. |
scramdb_vector_index_staleness | gauge | Writes accepted but not yet searchable. Reads 0: a write is searchable by the very next query, so there is nothing for this gauge to ever report as pending. |
scramdb_vector_index_compactions_total / scramdb_vector_index_compaction_running | counter / gauge | Compactions completed, and whether one is running now. |
scramdb_vector_index_fits_total | counter | Coarse model fits the index completed. |
scramdb_vector_index_maintenance_running / scramdb_vector_index_graph_armed | gauge | Whether the index's own maintenance is running, and whether its navigation graph is built and in use. |
scramdb_vector_index_edge_retries_total | counter | Navigation graph edge publications retried because another writer changed the same edge list first. A retry never loses the edge; a high rate means many threads link into the same nodes at once. |
scramdb_vector_index_unplaced_slots | gauge | Vectors the index's last model fit could not place in the partition it chose for them, for want of room there. While that is a material share of the index, the navigation graph stays off. |
scramdb_vector_index_io_reads_total / scramdb_vector_index_io_bytes_read_total | counter | Reads and bytes read for vectors spilled to disk. |
scramdb_vector_index_cache_hit_ratio | gauge | Share of vector reads served from memory. |
scramdb_vector_index_simd_level{level} | gauge | The instruction set the index runs with; the matching level reads 1. |
scramdb_vector_index_knob{knob} | gauge | The search knob value in force for the index. |
scramdb_vector_index_estimator_agreement | gauge | Share of candidates whose estimated order matched the exact order in recent searches. Emitted only once a real measurement exists. |
scramdb_vector_index_recall{k} | gauge | Recall measured by the last recall job for this index. Absent until a recall job has actually run, never a guessed value. |
scramdb_vector_index_last_build_wall_microseconds / scramdb_vector_index_last_build_cpu_microseconds | gauge | Wall time and CPU time (summed over every thread that worked on it) of the index's last build in this process. Absent until the index is built here: a restart reads absent until the next build, never a stale or guessed figure. |
scramdb_vector_index_last_build_peak_memory_bytes / scramdb_vector_index_last_build_rows | gauge | Resident memory of the index at the end of its last build (its peak: a build only adds), and rows it indexed. |
scramdb_vector_index_search_seconds | histogram | Time one search of the index took, end to end. |
scramdb_vector_index_last_flush_unix_seconds / scramdb_vector_index_rows_since_flush | gauge | When the index last made its writes durable, and rows written since then. |
scramdb_vector_index_last_reindex_unix_seconds | gauge | When the index was last rebuilt. Emitted only once a REINDEX has actually run. |
Cluster (labels index, table, group; present only once a distributed vector index
part exists):
| Metric | Type | Meaning |
|---|---|---|
scramdb_vector_group_owner_is_self | gauge | Whether this node owns the bucket group's index part. |
scramdb_vector_group_part_state{state} | gauge | The state of the index part for the bucket group (building, ready, regrowing, rebuilding, tail, invalid, absent, waiting_memory: a hosted bucket group this pass's own memory admission left out of the count it admitted to build). |
scramdb_vector_group_rebuilds_total / scramdb_vector_group_rebuild_running | counter / gauge | Times the index part was rebuilt after the bucket group moved, and whether one is running now. |
scramdb_vector_group_learner_ready | gauge | Whether a learner may serve vector searches for the bucket group. |
scramdb_vector_group_memory_in_use_bytes | gauge | Resident memory the index part for the bucket group uses. |
scramdb_vector_cluster_parts / scramdb_vector_cluster_parts_missing | gauge | Index parts this node holds for distributed tables, and bucket groups it owns whose part is not built. |
scramdb_vector_cluster_min_recall{index,k} | gauge | The lowest recall any of this node's parts of the index measured at its last recall run. Absent until measured; the whole cluster's lowest is the lowest over every node. |
Recall job (present only on a cluster node, where the job runs):
| Metric | Type | Meaning |
|---|---|---|
scramdb_vector_recall_parts_measured_total / scramdb_vector_recall_parts_skipped_total | counter | Index parts the recall job measured on this node, and parts a pass was due to measure and could not (each such part says why in scram_vector_index_parts.recall_note). |
scramdb_vector_recall_pass_cpu_seconds | histogram | CPU time of one recall pass on this node. |
scramdb_vector_recall_next_wait_seconds{reason} | gauge | Seconds until the job next measures a part, and why: churn (the wait follows how much of a part changed), cost (the one-percent CPU duty cycle binds), ceiling (a lightly changed part waits the hour ceiling), pinned (vector.recall_interval_secs), first (a part never measured), unchanged (nothing changed since the last measurement, so nothing is due: +Inf), or off. |
scramdb_vector_recall_probes{reason} | gauge | Probe queries the last measurement sampled, and why: precision (enough for a one-point standard error at the part's last recall), floor (8), ceiling (64), memory (the vector memory budget held fewer), or pinned (vector.recall_probes). |
Vector index system views
The same sample backs five read-only system views, so a dashboard that prefers SQL
over scraping Prometheus sees the identical numbers. Every row leads with node, the node
that measured it.
On a cluster, the four statistics views answer for the whole cluster from whichever node you
query: that node's rows and every other member's, gathered at query time. A member that does
not answer within vector.stats_node_timeout_ms (default 5000) contributes one row with its
node, the word unreachable in state (or metric), the reason in error, and NULL
everywhere else, never a silently shorter result. scram_vector_settings is the queried
node's own settings. On a single node every view is that node's rows.
| View | Columns |
|---|---|
scram_vector_settings | node, key, value, source, unit, live, description - every [vector] key, the value in force, and which layer set it |
scram_vector_indexes | node, index, schema, table, column, method, metric, opclass, state, arena_slots, arena_used_permille, regrows, live, tombstones, unindexed_rows, memory_budget, memory_in_use, disk_bytes, journal_bytes, search_qps, p50_us, p99_us, staleness, compactions, fits, graph_armed, simd_level, estimator_agreement, recall_at_10, last_flush, last_reindex, options, build_wall_us, build_cpu_us, build_peak_memory, build_rows, search_count, search_p99_us, error - one row per index per node; a distributed index's parts on a node fold into its row (build cost summed, NULL unless every part was built by that process; recall the lowest part's) |
scram_vector_index_parts | node, index, schema, table, group, owner, is_self, state, live, memory_in_use, rebuilds, learner_ready, rebuild_running, recall_at_10, recall_probes, recall_measured_at, recall_note, error - one row per bucket group each node hosts; is_self marks the rows of the node you queried; recall_note says why a part has no recall figure (too few vectors, over vector.recall_max_rows, not ready, or no memory); empty on a single node |
scram_vector_memory | node, total, reserved, in_use, engine_bytes, refusals, error - one summary row per node |
scram_vector_activity | node, metric, value, unit, description, error - the node-wide vector counters, one row per counter per node |
SELECT node, index, state, live, memory_in_use, search_qps FROM scram_vector_indexes;
SELECT node, "group", state, live, recall_at_10, recall_note
FROM scram_vector_index_parts WHERE index = 'docs_embedding_idx' OR error IS NOT NULL;
SELECT key, value, source, live FROM scram_vector_settings WHERE key = 'vector.over_fetch';
EXPLAIN on an indexed vector query names the tier chosen and the candidate count;
EXPLAIN ANALYZE adds candidates fetched, widening rounds taken, re-rank time, and
whether the answer came back underfilled. See Vector Search
for the query shapes these views and counters describe, and Tuning: Vector
Indexes for how to act on them.
What is not on this endpoint
- GPU metrics are not exposed today. There is no
scramdb_gpu_*family on/metrics, even though[gpu] metrics_enableddefaults totrue. See GPU Acceleration for the log-based way to confirm GPU dispatch instead. - No buffer-pool hit-ratio, cache-hit-ratio, or per-worker-utilization metric exists. Memory and parallelism verification on this site relies on
\timingand OS-level observation instead, see Memory and Parallelism.
Log levels and format
SCRAMDB_LOG=debug SCRAMDB_LOG_FORMAT=json scramdb -c scramdb-config.toml --pg-address 127.0.0.1:5432 --pg-no-auth
Expected: stderr lines shaped like {"ts":"...","level":"INFO","target":"...","msg":"Metrics endpoint listening on http://0.0.0.0:9090/metrics"}.
SCRAMDB_LOG
Controls the level filter. Accepted values (case-insensitive): off, error, warn, info, debug, trace. Default when unset: warn, so a release binary starts quiet. An unrecognized value does not fail loud, it silently falls back to warn, so a typo in this env var will not show up as an error, it just quietly serves you the default level.
SCRAMDB_LOG_FORMAT
Set to json (case-insensitive) for structured JSON output. Anything else, including leaving it unset, produces the default plain-text format.
Plain-text format: [{level}] [{timestamp}] {message}, for example:
[INFO] [2026-08-02T10:15:30.123Z] Metrics endpoint listening on http://0.0.0.0:9090/metrics
JSON format: {"ts":"<timestamp>","level":"<LEVEL>","target":"<module path>","msg":"<message>"}. Only backslash (\) and double-quote (") characters in the message are escaped. No other JSON escaping is performed, so a log message containing a raw newline or control character is not escaped and can produce an invalid JSON line in that rare case. Keep this in mind if you feed the JSON log output into a strict JSON-line parser downstream.
Logging is asynchronous: a dedicated low-priority background thread drains log messages and writes to stderr, so logging never blocks a query worker. Output always goes to stderr, never directly to a file, redirect it yourself (scramdb ... 2> scramdb.log) or capture it with your process supervisor.
Related diagnostic env vars
One additional env var is useful for deeper diagnosis alongside SCRAMDB_LOG, covered on another page in this section:
SCRAMDB_GPU_FORCE=1, forces GPU pipeline-shape eligibility for benchmarking. See GPU Acceleration.
Next
See Query Performance for how the JIT chunk counters are used in practice, GPU Acceleration for why GPU verification uses logs instead of this endpoint, or Clustering for how the cluster counters above map onto a Docker Compose or Kubernetes deployment.