Reference
Metrics
The 7 metric names
These names are contract-locked at runtime contract version 1. Renaming one is a breaking change with a version bump, not a refactor.
| Metric | Meaning |
|---|---|
trellara_runtime_live | 1 unless the phase is stopped or failed. |
trellara_runtime_ready | 1 only while the service is accepting work. |
trellara_runtime_pending_work | Work the service knows about and has not completed. |
trellara_runtime_restarts_total | Worker restarts since process start. |
trellara_runtime_failures_total | Failed worker attempts since process start. |
trellara_runtime_last_success_unixtime | When a durable boundary last completed. |
trellara_runtime_last_durable_lsn_bytes | The durable LSN as an integer, so it can be differenced. |
The label set, in full
Every metric carries exactly three labels and there is no configuration that adds a fourth.
| Label | Values |
|---|---|
service | relay · applier · iceberg_writer |
source_id | the configured source identity |
dataset_id | the configured dataset identity |
There is no label for a table, a transaction, a partition, an LSN or an error string. The label set is structurally incapable of the cardinality explosion that makes replication metrics expensive. Everything that needs that detail goes to trellara status, structured logs and evidence records.
Exposition
Served at /metrics as Prometheus text exposition 0.0.4. The endpoint is on --health-listen, which defaults to 0.0.0.0:9401 for the relay and 0.0.0.0:9402 for the applier.
What is deliberately not a metric
WAL headroom, quarantine contents and epoch state are not exported as metrics. They carry detail an operator needs to act, and that detail is exactly what makes a metric expensive. They come from commands instead — trellara check, trellara status --view diagnostics and trellara lake epoch — each of which returns JSON you can route.
Generated from the repository at
347a884 by tools/gen_reference.py. If this page and the
code disagree, the code is right and this page is a bug.