Operations
Production deployment
An evaluation config and a production config are not the same
document with different hostnames. environment: production
is a validation gate: set it and a class of configurations stops being accepted
at all.
Everything on this page is enforced in code today. What is not here is just as important: there is no tagged release to deploy, no upgrade or rollback path that has been tested, and no capacity model derived from a real workload. Those are named in the evidence table and they are the reason this page describes configuration rather than a rollout plan.
What production mode actually refuses
Configuration schema version 2 splits validation by environment. Under
environment: production, two rules are hard failures
rather than warnings:
- Every
database_url— source and target — must be a secret reference. An inline connection string is rejected. - A Kafka stream must carry
profile.kind: production. The development profile is refused outright.
This is deliberately blunt. A configuration that reaches production with an inline password is a configuration that will end up in a ticket, a chat message and a backup.
config_version: 2 environment: production source: id: retail-production database_url: source: environment_variable name: TRELLARA_SOURCE_DATABASE_URL publication: trellara_retail slot: trellara_retail_slot wal_retention_warn_bytes: 1073741824 pgoutput: protocol_version: 2 streaming: true
Secret references take one of two shapes, and both are validated when the config is parsed rather than when the connection is opened:
| Shape | Rule |
|---|---|
{ source: environment_variable, name: NAME } | NAME must be an ASCII shell identifier. |
{ source: file, path: /abs/path } | The path must be absolute. Relative paths are rejected. |
The shape of a flow config
The document is deny_unknown_fields: a typo in a key
is an error, not a silently ignored setting. There are four blocks.
| Block | Carries |
|---|---|
source | Identity, connection reference, publication, slot, capture kind, WAL warning threshold, spill threshold and directory, pgoutput protocol settings. |
dataset | Identity, mode, unknown-table policy, table list with verification keys, and optionally partition or strict_chunking. |
stream | kind: local with a path and durability, or kind: kafka with bootstrap servers, topic, consumer group and profile. |
target | Optional. A target connection reference and table contracts. Absent for a lake-only flow. |
There is no lake block. Lake planning derives its table
set from dataset.tables, which is what keeps the lake
output and the replication target describing the same thing.
The Kafka production contract
Version 1 of the contract is validated by
KafkaProductionContract::validate before a broker is
contacted. Every one of these is a rejection, not a warning:
| Requirement | Why it is not negotiable |
|---|---|
TLS, with mutual TLS or SASL (plain, scram_sha256, scram_sha512) | The stream carries committed row data. |
| CA certificate and credentials by reference | Same reason inline database URLs are refused. |
replication_factor ≥ 3 | The stream is the durability boundary the source acknowledgement depends on. |
min_insync_replicas ≥ 2, and never above the replication factor | A quorum of one is not a quorum. |
acks=all | A leader-only acknowledgement is not durable publication. |
| Producer idempotence enabled | Removes broker-side duplicates from retries. |
| Consumer auto-commit disabled | Auto-commit advances progress on a timer rather than on a successful apply. |
store_offset_after_apply and synchronous_offset_commit | The offset moves after the target transaction commits, or the ordering guarantee is gone. |
At runtime the production consumer validates every subscribed topic before joining the group, and the production publisher validates each topic before its first publish: the configured number of distinct brokers, and for every partition a healthy leader, distinct replicas and the configured live ISR quorum. A topic that is absent, under-replicated or below quorum is rejected before Trellara treats it as a durable boundary.
All of the above is enforced. None of it has been exercised against a broker under failure in CI. Leader loss, ISR shrink, rebalance and ambiguous delivery are covered by deterministic simulation, not by a live Kafka failure matrix.
Running the services
The relay and the applier are long-running processes with a shared lifecycle contract. Both take the same runtime flags:
trellara relay --config trellara.yml --health-listen 0.0.0.0:9401 --idle-poll-ms 250 --retry-initial-ms 250 --retry-max-ms 30000 --shutdown-grace-ms 30000 trellara apply --config trellara.yml --health-listen 0.0.0.0:9402
Reconnect backoff is capped at --retry-max-ms.
On SIGTERM the service drains: an in-flight durable
boundary gets up to --shutdown-grace-ms to finish, and
readiness goes false immediately so a load balancer stops sending work before the
process stops accepting it.
Both commands are hidden from trellara --help. They
work and are tested; they are not part of the nine-command public surface, which
is init, check,
preflight, run,
verify, status,
fleet, lake and
config.
Moving a config forward
trellara config migrate --config trellara.yml trellara config redact --config trellara.yml
migrate upgrades a known older schema to version 2
and refuses an unknown future version rather than guessing.
redact renders a structurally complete copy with secrets
removed — the version you attach to a ticket.
What this page deliberately does not tell you
How much WAL headroom you need, how many partitions to run, what the target lag will be, how long a reseed takes on your data, or how to roll from one version to the next. Every one of those requires a measurement that has not been taken. When they exist they will appear here with the run behind them, and not before.