Concepts
Snapshot and handoff
Initial copy is not a bootstrap side quest. It is the first correctness boundary in a verified flow: the moment where copied state has to meet streamed changes with neither a gap nor an unrecorded overlap between them.
What makes it hard
Copying a table takes time, and the source keeps committing while you copy. If CDC starts at the wrong LSN you either miss changes that happened during the copy or replay changes already contained in it. The bridge is a single value: the consistent LSN exported alongside the logical slot, which is both the visibility point the copy reads at and the position the stream starts from.
The state machine
Runs are keyed by source_id, dataset_id, run_id and
move through nine durable states. Backwards transitions are rejected; repeating
the same state is allowed so that idempotent retries can persist refreshed
counters.
| State | Meaning |
|---|---|
planned | The run exists but no slot or snapshot boundary has been established. |
slot_created | The logical slot exists and the run can acquire or reuse its consistency point. |
snapshot_exported | The source snapshot and consistent LSN are known. |
copying_table | At least one configured table is actively being copied. |
copy_complete | All required table copies reached the consistent snapshot boundary. |
stream_handoff_ready | Copy progress and handoff evidence are durable. The relay may now stream from the slot boundary. |
streaming | CDC has started after handoff. |
verified | Convergence verification promoted the run from handoff-ready to verified. |
failed_recoverable | A copy or handoff step failed after durable evidence exists. An operator can fix and resume. |
planned -> slot_created -> snapshot_exported -> copying_table -> copy_complete -> stream_handoff_ready -> streaming -> verified
failed_recoverable can only be entered once slot
boundary evidence exists, and must carry the slot name, the consistent LSN and a
failure reason so an operator knows which durable boundary can be retried. From
there a run can restart at planned,
slot_created, snapshot_exported
or copying_table. Moving from
verified back to failed_recoverable
is permitted only to record newly discovered evidence that a previous handoff was
unsafe.
The copy and handoff rules
- Create or reuse a
pgoutputlogical slot with an exported snapshot. - Record the slot name and consistent LSN before copying anything.
- Copy each configured table at the exported snapshot.
- Record table progress by
(source_id, dataset_id, run_id, relation). - Re-copying the same table for the same run is idempotent — the latest durable progress wins.
- Do not mark the run
stream_handoff_readyuntil every configured table is copied and handoff events are recorded at the same consistent LSN. - The relay may start CDC only after
stream_handoff_ready. - Verification may promote the run to
verifiedonly after target convergence succeeds.
A run with copied rows but no consistent LSN is not safe to stream. Copied data alone is not a boundary; the LSN is what makes it one. This is the failure mode that produces a target which looks complete, passes a row count, and is missing everything committed during the copy window.
Where the evidence lives
Three durable tables in the checkpoint schema, all of them customer-facing evidence rather than internal state:
| Table | Contents |
|---|---|
trellara.snapshot_runs | One row per run: state, slot name, consistent LSN, current relation, copied row count, failure reason, timestamps. |
trellara.snapshot_table_progress | One row per run and relation: relation state, copied row count, watermark LSN, update time. |
trellara.snapshot_handoff_events | Immutable handoff evidence: relation, copied rows, completion time, source watermark. |
Status output, reports, proof bundles and correctness checks read these rows. None of them infer handoff state from logs — which is why handoff state survives a log rotation, a restart, and a different operator picking the run up tomorrow.
Crash boundaries
| Failure | Required behaviour |
|---|---|
| Crash before slot creation | A later run starts from planned. |
| Crash after slot creation, before snapshot export | Retry records or reuses slot evidence before copying. |
| Crash after export, before copy | Retry starts table copy from the exported boundary. |
| Crash during table copy | Retry resumes or restarts that table without marking handoff ready. |
| Crash after all copies, before the handoff event | Retry records handoff evidence before streaming. |
| Crash after the handoff event, before the relay starts | The relay starts from the recorded slot boundary. |
| Duplicate snapshot attempt | Run and table primary keys make the operation idempotent. |
| DDL during copy | Contract evidence withholds handoff until the schema is refreshed. |
| Verification fails after copy | The run stays handoff-ready or recoverable — never verified. |