v0.1 pre-release · Apache-2.0 Start an evaluation

Operations

Monitoring and alerts

Operator · SRE The four endpoints, the seven metric names, and which of them should page someone.

Two surfaces, deliberately separated. Metrics answer “is this service healthy” with a label set small enough to be safe. Everything that needs a transaction id, an LSN or a table name goes to status, structured logs and evidence records, where cardinality does not matter.

The four endpoints

The relay and the applier each serve four paths on --health-listen, defaulting to 0.0.0.0:9401 and 0.0.0.0:9402.

PathReturnsUse it for
/livez200 unless the phase is stopped or failed; 503 otherwise.Restart policy. It answers “is this process worth keeping”.
/readyz200 only when the service is accepting work; 503 otherwise.Load-balancer and rollout gating. Draining reads 503 immediately.
/healthz200 with the full snapshot as JSON, always.Debugging and dashboards. Never as a probe — it is 200 even when blocked.
/metrics200, Prometheus text exposition 0.0.4.Scraping.
The distinction that matters

/livez and /readyz answer different questions and a common mistake is to wire both to the same probe. A service that is running but blocked — invalid slot, incomplete broker topology, unreconciled writer intent — is alive and not ready. Restarting it will not help; it needs an operator. Point your liveness probe at /livez or you will restart-loop a process that is correctly refusing to proceed.

The seven metric names

These names are contract-locked. Renaming one is a breaking change with a version bump, not a refactor.

MetricMeaning
trellara_runtime_live1 unless stopped or failed.
trellara_runtime_ready1 only while accepting work.
trellara_runtime_pending_workWork the service knows about and has not completed.
trellara_runtime_restarts_totalWorker restarts since process start.
trellara_runtime_failures_totalFailed worker attempts since process start.
trellara_runtime_last_success_unixtimeWhen a durable boundary last completed.
trellara_runtime_last_durable_lsn_bytesThe durable LSN as an integer, so it can be differenced.

Every one carries exactly three labels — service, source_id, dataset_id — where service is relay, applier or iceberg_writer. There is no label for a table, a transaction, a partition or an error string, and there is no configuration that adds one. The label set is structurally incapable of the cardinality explosion that makes replication metrics expensive.

Reading the lifecycle

/healthz reports a phase and a readiness, and only certain combinations are legal — an illegal one is a validation error, not a state you can observe.

PhaseReadinessAccepting work
starting, draining, stoppednot_readyno
failedblockedno
runningready or degradedyes
runningbackpressured or blockedno

degraded, backpressured and blocked always carry a reason_code — lowercase, digits and underscores, so it is safe to route on. That code is the difference between “something is wrong” and a runbook entry.

What should page someone

Trellara ships no alert rules, because a threshold without your workload behind it is a guess. What follows is the shape we would start from, and every number in it is yours to set.

ConditionWhySeverity
trellara_runtime_ready == 0 sustained, with the phase runningThe service is refusing work and a restart will not fix it. Read the reason code.Page
time() - trellara_runtime_last_success_unixtime above your freshness budgetNo durable boundary has completed. This is the closest thing to a real staleness signal.Page
rate(trellara_runtime_failures_total[5m]) rising while ready stays 1The service is retrying its way through something. Often the first sign of a source or broker problem.Ticket
Source WAL headroom falling, from trellara checkThe failure mode that damages the database rather than the pipeline. Alert on it before it is urgent.Page
Any quarantined transactionQuarantine is deliberate and never clears itself. A quarantined transaction that nobody looks at is a silent stall.Ticket
An epoch closing complete_with_gaps that your policy expected to be completeDownstream numbers are about to be published with a named gap.Ticket

WAL headroom, quarantine and epoch state are not metrics, deliberately. They come from commands that carry the detail an operator needs:

trellara status --config trellara.yml --view diagnostics --format json
trellara status --config trellara.yml --view alerts      --format json
trellara check  --config trellara.yml --format json

status takes --view of flow, report, alerts, dashboard, metrics or diagnostics. The JSON field names are a compatibility surface; the text rendering is for people and may change.

What is missing

No shipped dashboard, no alert-rule bundle, no recording rules, and no tracing. More importantly: no service-level objective anywhere on this page has a measurement behind it. Recovery has been proven correct in 63 deterministic scenarios and has never been timed, so any latency or recovery-time target you set today is your estimate rather than ours.

On this page