Concepts
Replay and deduplication
Trellara's transport is at-least-once and always will be. The guarantee that matters is not that a message arrives once, but that its effect happens once — and those are different properties with very different costs.
Why duplicates are the design
The checkpoint ordering deliberately publishes before acknowledging, which means every crash window that could have produced a gap produces a duplicate instead. Duplicates therefore are not an edge case to be minimised; they are the normal, expected output of every recovery path, and apply has to be built for them from the start rather than hardened against them later.
Sources of duplicate delivery, all of them routine:
- The relay crashed after publishing and before acknowledging the source; on restart it resumes from the last durable checkpoint and republishes.
- PostgreSQL redelivered from the slot because it never received feedback.
- The broker restarted and resolved an ambiguous acknowledgement by redelivering.
- The stream acknowledgement failed after a successful target apply.
- An operator replayed a quarantined transaction deliberately.
Transaction-level dedup
Deduplication is keyed at the transaction, not the row, and the dedup record is written inside the same target transaction as the data and the checkpoint:
BEGIN apply ordered changes for transaction T insert dedup record for T update target checkpoint to T.commit_lsn COMMIT
That single property is what makes the effect exactly-once. There is no window in which the data landed but the dedup record did not, because there is no sequence — they are one atomic write. A redelivery of T finds the record already present, skips the apply entirely, and acknowledges.
At-least-once delivery, exactly-once effect. Trellara never claims a message arrives once. It claims that applying it twice is indistinguishable from applying it once, because the second attempt does not apply.
Record-level keys, and what they are not for
Every ChangeRecord carries a deterministic
idempotency key:
{source_id}:{commit_lsn}:{transaction_id}:{total_order}
These are not what apply deduplicates on. They exist so that inspection, lake planning and future partial-replay analysis are deterministic — so that two independent tools examining the same change agree on what to call it. Building a consumer that deduplicates per-record instead of per-transaction is possible and will be slower, and it reintroduces the half-applied-transaction problem the protocol exists to prevent.
Duplicates under a manifest barrier
In chunked and partitioned modes a duplicate can arrive for an individual chunk rather than a whole transaction. This is safe for the same reason: a barrier-aware applier stages chunks, verifies the manifest, reconstructs the transaction, and then performs one atomic apply that carries the transaction-level dedup record. Duplicate chunk replay changes what is staged; it cannot change what is applied.
Confirming it afterwards
The claim is falsifiable, which is the point. Kill the relay, the applier or the broker at any point in a flow, let it recover, then compare source and target directly:
trellara verify --config trellara.yml
Row counts, primary-key checksums and watermarks are compared for real rather than inferred from offsets. A replayed duplicate that had been double-applied would show up as a row-count or checksum divergence, not as a silent success. See the verification model for what that command does and does not check.