v0.1 pre-release · Apache-2.0 Start an evaluation

Concepts

Snapshot and handoff

Developer · Operator The first correctness boundary in a flow, and the one most often treated as a bootstrap detail.

Initial copy is not a bootstrap side quest. It is the first correctness boundary in a verified flow: the moment where copied state has to meet streamed changes with neither a gap nor an unrecorded overlap between them.

What makes it hard

Copying a table takes time, and the source keeps committing while you copy. If CDC starts at the wrong LSN you either miss changes that happened during the copy or replay changes already contained in it. The bridge is a single value: the consistent LSN exported alongside the logical slot, which is both the visibility point the copy reads at and the position the stream starts from.

The state machine

Runs are keyed by source_id, dataset_id, run_id and move through nine durable states. Backwards transitions are rejected; repeating the same state is allowed so that idempotent retries can persist refreshed counters.

StateMeaning
plannedThe run exists but no slot or snapshot boundary has been established.
slot_createdThe logical slot exists and the run can acquire or reuse its consistency point.
snapshot_exportedThe source snapshot and consistent LSN are known.
copying_tableAt least one configured table is actively being copied.
copy_completeAll required table copies reached the consistent snapshot boundary.
stream_handoff_readyCopy progress and handoff evidence are durable. The relay may now stream from the slot boundary.
streamingCDC has started after handoff.
verifiedConvergence verification promoted the run from handoff-ready to verified.
failed_recoverableA copy or handoff step failed after durable evidence exists. An operator can fix and resume.
planned -> slot_created -> snapshot_exported -> copying_table
  -> copy_complete -> stream_handoff_ready -> streaming -> verified

failed_recoverable can only be entered once slot boundary evidence exists, and must carry the slot name, the consistent LSN and a failure reason so an operator knows which durable boundary can be retried. From there a run can restart at planned, slot_created, snapshot_exported or copying_table. Moving from verified back to failed_recoverable is permitted only to record newly discovered evidence that a previous handoff was unsafe.

The copy and handoff rules

  1. Create or reuse a pgoutput logical slot with an exported snapshot.
  2. Record the slot name and consistent LSN before copying anything.
  3. Copy each configured table at the exported snapshot.
  4. Record table progress by (source_id, dataset_id, run_id, relation).
  5. Re-copying the same table for the same run is idempotent — the latest durable progress wins.
  6. Do not mark the run stream_handoff_ready until every configured table is copied and handoff events are recorded at the same consistent LSN.
  7. The relay may start CDC only after stream_handoff_ready.
  8. Verification may promote the run to verified only after target convergence succeeds.
The rule that catches people

A run with copied rows but no consistent LSN is not safe to stream. Copied data alone is not a boundary; the LSN is what makes it one. This is the failure mode that produces a target which looks complete, passes a row count, and is missing everything committed during the copy window.

Where the evidence lives

Three durable tables in the checkpoint schema, all of them customer-facing evidence rather than internal state:

TableContents
trellara.snapshot_runsOne row per run: state, slot name, consistent LSN, current relation, copied row count, failure reason, timestamps.
trellara.snapshot_table_progressOne row per run and relation: relation state, copied row count, watermark LSN, update time.
trellara.snapshot_handoff_eventsImmutable handoff evidence: relation, copied rows, completion time, source watermark.

Status output, reports, proof bundles and correctness checks read these rows. None of them infer handoff state from logs — which is why handoff state survives a log rotation, a restart, and a different operator picking the run up tomorrow.

Crash boundaries

FailureRequired behaviour
Crash before slot creationA later run starts from planned.
Crash after slot creation, before snapshot exportRetry records or reuses slot evidence before copying.
Crash after export, before copyRetry starts table copy from the exported boundary.
Crash during table copyRetry resumes or restarts that table without marking handoff ready.
Crash after all copies, before the handoff eventRetry records handoff evidence before streaming.
Crash after the handoff event, before the relay startsThe relay starts from the recorded slot boundary.
Duplicate snapshot attemptRun and table primary keys make the operation idempotent.
DDL during copyContract evidence withholds handoff until the schema is refreshed.
Verification fails after copyThe run stays handoff-ready or recoverable — never verified.

On this page