v0.1 pre-release · Apache-2.0 Start an evaluation

Concepts

Checkpoints and acknowledgements

Developer · Operator The ordering rule that decides whether a crash costs you a duplicate or a gap.

Almost every durability property Trellara claims reduces to one ordering decision: what is made durable before what is acknowledged. Get it backwards and a crash in a narrow window loses committed transactions permanently.

Three watermarks, not one

“Lag” is a single number pretending to be three. Trellara tracks them separately because they fail independently.

WatermarkMeaning
last_seen_lsnThe highest source LSN observed. Says nothing about durability.
last_durable_lsnThe highest source LSN durably published. This is what may be acknowledged to the source.
last_applied_lsnThe highest source LSN atomically applied at the target.

A gap between last_seen_lsn and last_durable_lsn is normal in-flight work. A gap between last_durable_lsn and last_applied_lsn is a consumer falling behind — and it is Trellara's retention problem at that point, not your primary's.

The relay invariant

publish durable message(s)
  -> save source durable checkpoint
  -> acknowledge source LSN

The relay must not advance the source durable checkpoint until the envelope — or every chunk plus its manifest and commit marker — has been durably published. Reversing these two steps makes the happy path marginally cheaper and loses committed transactions on exactly the failure below.

The applier invariant

apply target changes
  + record transaction dedup
  + save target checkpoint

These three effects share one target transaction. That is the whole mechanism. If apply fails, neither the checkpoint nor the dedup record advances, because they roll back with the data. If stream acknowledgement fails after a successful apply, redelivery arrives, the dedup record is already there, the transaction is skipped, and the acknowledgement is retried.

The trade being made

Publishing before acknowledging means a crash in the window between them produces a duplicate. Acknowledging before publishing would mean the same crash produces a gap. Duplicates are absorbable by a transactional dedup record; gaps are not recoverable from anywhere. Trellara chooses duplicates, everywhere, without exception.

Failure windows and required behaviour

Failure windowRequired behaviour
Relay exits before publishThe source durable checkpoint does not advance. Nothing was acknowledged; nothing is lost.
Publish accepted but the ack is lostA duplicate publish is allowed. No transaction may be missed.
Relay exits before source feedbackThe source may redeliver. Downstream dedup absorbs the duplicates.
Target apply fails before commitNo checkpoint or dedup advancement — they were in the rolled-back transaction.
Target checkpoint write fails during applyThe whole apply transaction rolls back.
Stream ack fails after applyRedelivery is skipped by transaction dedup, then acknowledged.
A manifest chunk is missingThe reconstructed transaction is not applied.

Cancellation

Shutdown is not an exception to any of this. The runtime's shutdown order is fixed:

stop admitting new work
  -> mark draining / not_ready
  -> complete or preserve in-flight durable boundaries
  -> flush checkpoint and receipt writes
  -> stop workers
  -> mark stopped / not_ready

The governing rule: cancellation must never convert an unproven publish, apply, upload or catalog outcome into success. A SIGTERM during an ambiguous operation resolves to “unknown, retry” and never to “done.”

When the transport is Kafka

The same invariant has to be expressed in the broker's vocabulary, and the production contract enforces it in config validation rather than documenting it and hoping:

A single-broker development configuration stays valid for local testing but does not satisfy the production contract and is never reported as production-ready.

On this page