Concepts
Transaction boundaries
An order, its line items, the payment and the inventory decrement were committed together because they mean nothing apart. Trellara's primary unit of correctness is one committed PostgreSQL transaction, and every other decision in the protocol follows from refusing to break it.
The envelope
Capture assembles a committed transaction into a TransactionEnvelope
— a versioned protobuf container that travels intact through transport to apply.
The fields that carry correctness:
| Field | Meaning |
|---|---|
protocol_version | Currently 1. A consumer must reject an envelope whose version it does not understand. |
source_id / database_id / dataset_id | Stable source, database and flow identity. |
transaction_id | Source transaction identity chosen by capture. |
begin_lsn / commit_lsn | Transaction begin and commit LSN. commit_lsn is the durable ordering and checkpoint boundary. |
commit_timestamp_ms | Source commit timestamp. |
schema_versions | Relation schema version fingerprints at capture time. |
changes | Ordered ChangeRecord list. |
ddl_events | DDL metadata sharing the same transaction boundary. |
checksum | xxh3 over the canonical envelope with the checksum field zeroed. |
Change records
Each ChangeRecord is one table operation inside the
transaction, and carries three separate orderings rather than one:
total_order across the whole transaction,
table_order within a relation, and
partition_order within a routed partition. Keeping all
three means a consumer can reconstruct source order after any reshuffling that
transport imposed.
Every record also carries a deterministic idempotency key:
{source_id}:{commit_lsn}:{transaction_id}:{total_order}
Appliers use transaction-level deduplication for exactly-once target effects. The record-level keys exist so that inspection, lake planning and future partial-replay analysis are deterministic — not because apply needs them.
Row images and replica identity
Protocol version 1 supports ordinary primary-key tables on
REPLICA IDENTITY DEFAULT. Two rules matter and both are
fail-safe in one direction:
- When
pgoutputomits an unchanged TOAST column, an applier must treat the absent non-key column as unchanged, not overwrite it with null. Getting this wrong silently destroys data. - A missing key column on an update or delete is fail-closed. There is no best-effort path.
Capture protocol scope
The default production path is pgoutput protocol
version 2 with streaming enabled:
source: pgoutput: protocol_version: 2 streaming: true
Version 2 lets Trellara consume PostgreSQL's streamed in-progress transactions
from START_REPLICATION while still imposing its own
manifest and commit-barrier semantics downstream. Version 1 is accepted for older
or constrained environments but is scoped to streaming: false,
which means a large transaction must be fully assembled before Trellara can chunk
it.
Three boundary modes, two configuration values
This trips people up, so it is worth stating before the modes themselves.
dataset.mode takes exactly two values —
strict_transaction_order and
partitioned_scale_mode. There is no
strict_chunked mode value. Chunking is a sub-option of strict mode,
switched on for the whole flow by dataset.strict_chunking.max_changes_per_chunk,
and the choice is made once per flow rather than per transaction. On the wire, the manifest
carries a separate enum, ManifestBoundaryMode, whose values are
strict_chunked_transaction_order and
partitioned_scale_mode. Operator output reports the wire vocabulary,
which is why status can name a chunked boundary your configuration
never spells that way.
Strict transaction order
One envelope per committed source transaction, published to the strict topic. Source order is preserved by commit LSN, target apply is atomic per transaction, and the target checkpoint and dedup record are written in the same target transaction. Stream acknowledgement happens only after a successful apply or a duplicate skip. This is the default for the brokerless local path and for ordinary Postgres-to-Postgres replication.
Strict chunked transaction order
For flows whose transactions can grow large enough that holding one envelope in memory is
unreasonable. Once strict_chunking is configured the relay emits chunks
on the strict topic, then a TransactionManifest, then a
TransactionCommitMarker — for every transaction on the flow,
including a one-change transaction. The transaction becomes visible only after a barrier-aware
applier has reconstructed every manifest-listed chunk and verified the manifest. Setting the
threshold is therefore a decision about the flow's worst case, not a switch that only fires on
large transactions.
The chunked mode reuses the PartitionChunk payload
type, so a chunk carries a partition_id. In this mode
that value is a chunk index, not an ownership partition. Consumers that
conflate the two will route correctly by accident and incorrectly under load.
Partitioned scale mode
Changes are routed by a configured ownership key, one chunk per participating partition, then a manifest and a commit marker. A partition-local consumer may process its own lane for low-latency local work, but it must treat the transaction as incomplete without the barrier. Global current-state visibility, Postgres apply and correctness reports all require the complete manifest plus commit marker.
Partition policies
Ownership keys are not always present or stable, so both cases are explicit policy rather than implementation accident.
| Policy | Behaviour |
|---|---|
null_key_policy = quarantine | Fail closed when the ownership key is null. |
route_to_singleton_partition | Route null ownership keys to partition 0. |
route_to_dead_letter_partition | Route null ownership keys to the final partition so local consumers can isolate them. |
derive_from_primary_key | Route by hashing stable primary-key columns. |
key_change_policy = quarantine | Fail closed when an update moves ownership between partitions. |
key_change_policy = emit_move | Emit a delete and an insert under one transaction manifest. |
dual_write_window and forbid
exist as declared policy names for contract documentation, but the partition
planner fails closed unless ownership moves are handled by
emit_move.
What fails closed
Every one of these is a correctness failure rather than a warning, and every one of them stops apply:
- A missing chunk from a manifest-listed set.
- A transaction id mismatch between a chunk and its manifest.
- An event-count mismatch against
global_event_count. - Any checksum mismatch — envelope, chunk, manifest, or commit marker.
- An empty manifest barrier. A relay must not publish a manifest and commit marker with zero chunks or zero source changes; protocol planning rejects it.
- A duplicate
total_orderbetween a row change and a DDL event, which would make target replay ambiguous.