Operations
Monitoring and alerts
Two surfaces, deliberately separated. Metrics answer “is this
service healthy” with a label set small enough to be safe. Everything that
needs a transaction id, an LSN or a table name goes to
status, structured logs and evidence records, where
cardinality does not matter.
The four endpoints
The relay and the applier each serve four paths on
--health-listen, defaulting to
0.0.0.0:9401 and 0.0.0.0:9402.
| Path | Returns | Use it for |
|---|---|---|
/livez | 200 unless the phase is stopped or failed; 503 otherwise. | Restart policy. It answers “is this process worth keeping”. |
/readyz | 200 only when the service is accepting work; 503 otherwise. | Load-balancer and rollout gating. Draining reads 503 immediately. |
/healthz | 200 with the full snapshot as JSON, always. | Debugging and dashboards. Never as a probe — it is 200 even when blocked. |
/metrics | 200, Prometheus text exposition 0.0.4. | Scraping. |
/livez and /readyz answer
different questions and a common mistake is to wire both to the same probe. A
service that is running but
blocked — invalid slot, incomplete broker topology,
unreconciled writer intent — is alive and not ready.
Restarting it will not help; it needs an operator. Point your liveness probe at
/livez or you will restart-loop a process that is
correctly refusing to proceed.
The seven metric names
These names are contract-locked. Renaming one is a breaking change with a version bump, not a refactor.
| Metric | Meaning |
|---|---|
trellara_runtime_live | 1 unless stopped or failed. |
trellara_runtime_ready | 1 only while accepting work. |
trellara_runtime_pending_work | Work the service knows about and has not completed. |
trellara_runtime_restarts_total | Worker restarts since process start. |
trellara_runtime_failures_total | Failed worker attempts since process start. |
trellara_runtime_last_success_unixtime | When a durable boundary last completed. |
trellara_runtime_last_durable_lsn_bytes | The durable LSN as an integer, so it can be differenced. |
Every one carries exactly three labels —
service, source_id,
dataset_id — where service
is relay, applier or
iceberg_writer. There is no label for a table, a
transaction, a partition or an error string, and there is no configuration that
adds one. The label set is structurally incapable of the cardinality explosion
that makes replication metrics expensive.
Reading the lifecycle
/healthz reports a phase and a readiness, and only
certain combinations are legal — an illegal one is a validation error, not a
state you can observe.
| Phase | Readiness | Accepting work |
|---|---|---|
starting, draining, stopped | not_ready | no |
failed | blocked | no |
running | ready or degraded | yes |
running | backpressured or blocked | no |
degraded, backpressured and
blocked always carry a
reason_code — lowercase, digits and underscores, so
it is safe to route on. That code is the difference between “something is
wrong” and a runbook entry.
What should page someone
Trellara ships no alert rules, because a threshold without your workload behind it is a guess. What follows is the shape we would start from, and every number in it is yours to set.
| Condition | Why | Severity |
|---|---|---|
trellara_runtime_ready == 0 sustained, with the phase running | The service is refusing work and a restart will not fix it. Read the reason code. | Page |
time() - trellara_runtime_last_success_unixtime above your freshness budget | No durable boundary has completed. This is the closest thing to a real staleness signal. | Page |
rate(trellara_runtime_failures_total[5m]) rising while ready stays 1 | The service is retrying its way through something. Often the first sign of a source or broker problem. | Ticket |
Source WAL headroom falling, from trellara check | The failure mode that damages the database rather than the pipeline. Alert on it before it is urgent. | Page |
| Any quarantined transaction | Quarantine is deliberate and never clears itself. A quarantined transaction that nobody looks at is a silent stall. | Ticket |
An epoch closing complete_with_gaps that your policy expected to be complete | Downstream numbers are about to be published with a named gap. | Ticket |
WAL headroom, quarantine and epoch state are not metrics, deliberately. They come from commands that carry the detail an operator needs:
trellara status --config trellara.yml --view diagnostics --format json trellara status --config trellara.yml --view alerts --format json trellara check --config trellara.yml --format json
status takes --view of
flow, report,
alerts, dashboard,
metrics or diagnostics. The
JSON field names are a compatibility surface; the text rendering is for people and
may change.
What is missing
No shipped dashboard, no alert-rule bundle, no recording rules, and no tracing. More importantly: no service-level objective anywhere on this page has a measurement behind it. Recovery has been proven correct in 63 deterministic scenarios and has never been timed, so any latency or recovery-time target you set today is your estimate rather than ours.