v0.1 pre-release · Apache-2.0 Start an evaluation

Verified change data infrastructure for PostgreSQL

Every transaction.
Every database.
Provably accounted for.

Trellara is the operating layer above PostgreSQL logical replication. It keeps a stalled consumer from threatening your primary, moves committed transactions to Postgres targets and your lakehouse with their boundaries intact, and answers “did everything converge?” as a command that returns evidence — not a dashboard that returns a feeling.

Apache-2.0 Nothing installed in your database No broker required to start Your data never leaves your network RDS · Aurora · Cloud SQL · Azure Flexible Server · self-hosted
read-only · no state written
trellara-check $POSTGRES_URL --format text grade B capture readiness 4 tables ready · 1 note (TOAST) slots 3 active · 1 stalled 2h14m wal headroom ~36h at current lag failover slots supported · not enabled actions 2 recommended nothing was created — no publication, slot, or checkpoint trellara chaos run ✓ relay crash before publish nothing lost ✓ relay crash after publish duplicates only ✓ applier crash mid-transaction no double-apply ✓ partition barrier incomplete held, not exposed ✓ pinned schema drift capture stopped closed … 58 more scenarios trellara run --config trellara.yml --local --verify transport local segment log · fsync transactions 1,204 applied atomically CONVERGED

Command output above is illustrative — the command names, flags and output shapes are real; the values are an example fleet, not a capture from a customer system. We have no customer systems yet. What we have actually established →

Every claim on this page carries a state

Shipping

Runs today in the open-source binary. You can execute the named command against your own database this afternoon.

In progress

Contract is locked and code exists; the production path is narrower than the vision. Scope stated honestly in each section.

Designed

Specified in the public design docs, not built. Listed so you can judge direction, priced at zero in an evaluation.

Research

Explored and deliberately deprioritised. Named because pretending otherwise is how roadmaps lose credibility.

01 · The gap

PostgreSQL gives you replication primitives. It does not give you an operating model.

Native logical replication keeps getting better — failover slots, parallel apply, idle-slot timeouts, sequence sync. None of that changes three facts: a consumer's health is still coupled to your primary's disk, nothing tells you the target actually converged, and at fifty databases you are operating fifty independent pipelines by hand. We wrote the full comparison — including where native wins.

Without — slot pressure on the primary
21:04

A consumer hits a bad record and silently stops advancing its slot.

23:30

The slot retains WAL. Disk climbs. No alert fires — lag reads as “catching up”.

02:47

Storage alarms. On-call learns what a replication slot is during the incident.

03:15

Slot dropped, or invalidated by max_slot_wal_keep_size. Full resync.

With Trellara — pressure moved to a durable boundary
21:04

The consumer stalls against Trellara's durable log — not your primary's WAL.

21:06

trellara-check grades it: stalled consumer, WAL headroom in hours, two named actions.

21:20

The bad transaction sits in quarantine with a replay command. Unrelated flows keep moving.

09:00

Replay, then verify. Converged, with checksum evidence.

02 · What Trellara claims

Four claims. Each one is falsifiable by a command you run against your own system.

Everything else this project does is in service of these four. Each claim states its mechanism, names the command that would disprove it, and carries its maturity. The capabilities underneath are there if you want them — but if you only read four things, read these.

CLAIM 01

Your source database and its WAL stay protected.

Shipping

A logical replication slot is a standing claim on your primary's disk. When a consumer stalls, the slot holds WAL and the pressure lands on the database that is serving your customers. Postgres will not tell you this is happening; lag simply reads as “catching up” until storage alarms.

Trellara moves that pressure off the primary. The relay publishes to a durable boundary — a local segment log or your broker — and only then advances the source checkpoint. A consumer that stalls stalls against Trellara's retention, not yours. And the first command you ever run is a read-only grader that tells you the risk you already have, before Trellara creates anything.

How you falsify it

Point the diagnostic at a production source. It creates no publication, no slot, no checkpoint and no target state — verify that with your own audit log.

trellara-check $POSTGRES_URL --format text

Then stall a consumer on purpose and watch which disk grows. How the ordering works →

The capabilities underneath
Source safety & WAL governance Shipping

Grades an existing source for CDC readiness before anything is created: slot posture, plugin, wal_status, safe WAL size, invalidation reason, failover and synced flags, idle-slot timeout, and replication conflicts.

trellara-check · trellara preflight

Publish-before-acknowledge ordering Shipping

The source checkpoint advances only after the publish is durable. The failure mode this chooses is duplicates on replay, never a gap — and the applier is built to absorb duplicates.

see the recovery sequence in §03

Durable log without a broker Shipping

Disk-backed segment log with fsync, offset index, per-consumer cursors, torn-tail detection and large-transaction spill. Retention pressure lives here instead of on your primary, and the first evaluation needs no infrastructure.

trellara run --local

One slot per source Shipping

Replicas, search indexes, caches and lakehouse tables are served from a single capture. Every additional consumer costs a cursor on Trellara's log, not another slot on your database.

trellara status --view diagnostics

CLAIM 02

What Postgres committed as one fact arrives as one fact.

Shipping

An order, its line items, the payment and the inventory decrement were committed together because they mean nothing apart. Most change-data pipelines dissolve that transaction into a stream of independent row events and leave reassembly to whoever reads it — which in practice means nobody reassembles it, and every consumer occasionally sees a half-applied business fact.

Trellara treats the transaction boundary as part of the protocol, not as metadata. Committed transactions are assembled into a versioned envelope carrying commit LSN, schema fingerprint, row sequence and checksum. Strict mode preserves source commit order; partitioned mode preserves per-key order and reassembles cross-partition transactions behind manifest barriers before a reader can see any of it.

How you falsify it

Commit a multi-table transaction at the source, then read the target mid-flight and try to observe it half-applied.

trellara semantics --config trellara.yml

semantics is hidden from the short help surface, like chaos and pilot-* — the CLI reference lists every hidden command. The strict-chunk simulation family exists to attempt exactly this, five ways, on every change. The boundary model →

The capabilities underneath
Transaction boundaries as protocol Shipping

The envelope is an open, versioned contract carried as a protobuf IDL in the repository — protocol version, boundaries, commit LSN, schema fingerprint, ordered row sequence, DDL events, and an xxh3_64 checksum over the whole message. Thirteen fields, fixed tag numbers, validated on both encode and decode. Never optional, never a side topic, never a field a sink may ignore.

docs/DESIGN.md

Schema contracts & DDL barriers Shipping

Live pgoutput relation metadata is fingerprinted and pinned. Eleven change kinds are classified on three independent axes — compatibility (compatible · requires-mapping · destructive-or-ambiguous · blocked-by-policy), the resulting decision (auto-apply · stage-then-apply · manual-review · block), and the apply mode you asked for. Only three kinds are ever compatible — add_nullable_column, widen_type, increase_varchar — and only on a table already in your config. add_table joins them only under unknown_table_policy: allow_compatible. Renames require an explicit mapping. Drops, type narrowing, NOT NULL additions and primary- or partition-key changes are blocked whatever mode you pass. Each emits a propagation contract: barrier id, affected sinks, ordered phases, per-sink acknowledgement, and the visibility pause before post-DDL rows are released.

concept → · trellara schema ddl-plan

Atomic apply at the target Shipping

Each source transaction is applied atomically. Data and checkpoint commit together in one target transaction, so a crash mid-apply rolls the whole thing back rather than leaving the checkpoint ahead of the data.

trellara run --config trellara.yml

Cross-partition manifest barriers Shipping

Where throughput requires partitioning, a transaction spanning partitions is held behind a manifest barrier until every chunk has landed. Held, never half-exposed — including when the manifest arrives before the chunks do.

5 deterministic strict-chunk scenarios

CLAIM 03

You can recover from failure and prove you converged afterwards.

Shipping

Most replication tools tell you about lag. Lag is a statement about how far behind you are, not about whether what arrived is correct. After an incident, a migration or a reseed, the only question that matters is whether the destination actually matches the source — and almost nothing in this category will answer it.

Trellara answers it as a command that returns evidence: row counts, primary-key checksums and source-to-target watermarks, for a whole flow or one table. Every failure state names its own recovery command, incompatible data is quarantined rather than silently applied, and the whole failure surface is rehearsed deterministically on every change.

How you falsify it

Kill the relay, the applier or the broker mid-flow — at any point you like — then run:

trellara verify --config trellara.yml

It returns CONVERGED with checksums, or it names the divergence and the recovery path. There is no third answer. What verify does and does not check →

The capabilities underneath
Convergence verification Shipping

Row counts, primary-key checksums and source-to-target watermarks compared on demand, for a whole flow or one table at a time. After an incident you check rather than assume.

trellara verify --config trellara.yml --table public.sales

Quarantine, replay & reseed Shipping

Incompatible data is quarantined rather than silently applied or allowed to block unrelated flows. A destination past retention is re-snapshotted with an explicit visibility boundary and a verified handoff back to the live stream.

trellara repair-plan · reseed · stream locate-local

Deterministic failure matrix Shipping

63 scenarios across 28 deterministic simulations, each naming its protected invariant, its proof command and its recovery command. Runs in CI on every change and ships in the binary so you can run it against your own configuration.

trellara chaos run · trellara chaos report

Production runtime posture Shipping

Relay and applier are long-running services with /livez, /readyz, /healthz and Prometheus /metrics; capped exponential backoff; SIGTERM drain that preserves the in-flight durable boundary; SIGHUP config reload.

relay :9401 · applier :9402

CLAIM 04

Completeness across the whole fleet becomes a query, not an assumption.

In progress

At fleet scale the question changes. It stops being “is this pipeline healthy” and becomes “is last night's revenue number complete, and if not, which databases are missing from it?” Store databases lose connectivity, tenants migrate, regions fail over. Every tool in this category reports freshness; none of them report completeness.

Trellara fans every source's transaction stream into an append-only Iceberg changelog, and every commit lands alongside an epoch record: which databases are included, through which commit LSN, who is lagging, who is quarantined. Your revenue job joins against _trellara_epochs instead of guessing. The gap is named rather than averaged away.

Scope today, stated plainly: fleet config, reporting and scorecards ship; a continuously running multi-source fleet runtime is the next milestone. The Iceberg writer is deliberately narrow — append-only, REST catalog, S3-compatible, no native deletes. Current-state and SCD2 remain Spark templates by design.

How you falsify it

Take a source offline, let an epoch close, and check whether the epoch record admits the gap or quietly rounds over it.

trellara lake epoch --config fleet.yml trellara fleet scorecard --config fleet.yml

A closed epoch missing a source must read complete_with_gaps, never complete — and a consumer only sees it at all if it opted in, because the gate rejects gapped epochs by default. The epoch model →

The capabilities underneath
Fleet operations In progress

Topology, convergence gates, recovery drills and proof commands for many databases as one reviewable object. Available today as fleet config, reporting and scorecards.

trellara fleet init · report · scorecard

Lakehouse fan-in with completeness In progress

Per-source watermarks and an epoch record on every Iceberg commit, deduplicated across reconnects and batched through one committer per table so there is no small-file storm and no commit contention.

trellara lake plan · epoch · fanin verify

Offline-tolerant sources In progress

Nodes that go dark buffer locally and catch up in order when they return. A node past retention is reseeded and re-verified rather than silently diverging.

4 deterministic fleet fan-in scenarios

Tenant-boundary attestation Designed

For platforms running a database per customer: a policy layer that proves which per-tenant data crossed into a shared view and which did not. Specified in the design docs, not built. Priced at zero today.

analyzed — docs/ROADMAP.md

03 · Evidence

What we have established, and what we have not.

This is the shape of an enterprise evidence pack with a great many of its rows still empty. We publish it that way because the empty rows are the accurate answer, and because you would find them in week two of an evaluation anyway.

Established

Countable in the repository or produced by CI today. Every one of these is something you can recount from the source yourself.

Latest commit exercised by CI
347a884 · 2026-09-11Every push to main and every pull request runs format, clippy with warnings denied, the full workspace test suite, three release builds, example-config validation, the no-broker quickstart proof, a compose-config check and a correctness-report freshness gate.
Deterministic failure scenarios
63Across 28 simulations in five failure-point families. Each names its protected invariant, its proof command and its recovery command.
Correctness report
rebuilt daily · 09:17 UTCGenerated from the simulations on a schedule and deployed to GitHub Pages, where you can read it right now. CI fails if the committed report has drifted from the code.
Source protocol
pgoutput v2, streamingStandard PostgreSQL logical replication, and the default path. protocol_version: 2 with streaming is the default; v1 is accepted only with streaming disabled. On this path there is no extension, no superuser-owned agent and nothing installed in the database. The optional native extension is a separate, opt-in path — see below.
Live PostgreSQL in CI — native extension
15, 16, 17, 18The extension workflow runs a four-major matrix. Each job builds against real server headers, runs initdb, installs the extension, exercises every implemented decoding callback, asserts unchanged-TOAST markers survive an unrelated update, injects a relay failure and asserts the logical slot does not advance without a durable acknowledgement, then restarts the background worker and drains the replayed slot. It also builds Debian and RPM packages and verifies their contents and checksums.
Live PostgreSQL in CI — external relay
noneThe default path has no CI Postgres lane at all. The compose harness runs postgres:16 and that is the only version the external relay has been run against, by hand. Read the row opposite before you read either of these as a support matrix.
Optional native extension
Deb + RPM, PG 15–18A bounded shared-memory queue (16 slots, 64 KiB per frame) hands committed frames to a local relay over a mode-0600 Unix socket. Broker and network I/O stay outside the server process. Two-phase commit is unsupported and every prepared-transaction callback fails closed rather than dropping changes.
Kafka transport
librdkafka via rdkafka 0.37Exercised against Redpanda. Production contract v1 is enforced by KafkaProductionContract::validate, not documented and hoped for: TLS with mTLS or SASL SCRAM, credentials by reference only, RF ≥ 3, min.insync.replicas ≥ 2 and never above RF, acks=all, producer idempotence, consumer auto-commit off, and synchronous offset commit only after apply. A config that violates any of these is rejected before a broker is contacted.
Iceberg output
REST catalog · S3-compatibleAppend-only Parquet under a catalog you own. No native deletes; current-state and SCD2 are Spark templates.
Production config enforcement
schema v2, validatedenvironment: production is a gate, not a label. It refuses inline database URLs — every credential must be an environment-variable or absolute-file reference — and refuses any Kafka profile other than the production contract. config redact renders a shareable copy; secrets carry a Debug impl that prints <redacted>.
Observability surface
4 endpoints · 7 metrics/livez, /readyz, /healthz, /metrics. Seven contract-locked metric names, each carrying exactly three labels — service, source_id, dataset_id. The label set is structurally incapable of per-transaction or per-table cardinality.
Release artifacts
3 binaries · SHA-256Linux x86_64: trellara-check, default trellara, trellara-full. Checksums published with the release.
Toolchain
Rust 1.97.1, pinnedPinned in rust-toolchain.toml and matched in CI, so a build is reproducible from the repository alone.
Licence
Apache-2.0The wire format is a protobuf IDL in the repository, not a private encoding; the durable log format is documented frame by frame; Iceberg output is standard Parquet under a catalog you own. If you stop using Trellara your data stays readable and your consumers keep working.

Not yet established

Things an enterprise evaluation will ask for that we cannot answer today. Each is a real gap, not a roadmap item in disguise.

Tagged release
noneThe version string reads v0.1 pre-release. There is no v0.1.0 tag yet, so there is nothing to pin a dependency to.
PostgreSQL matrix for the default path
not establishedThe native extension has a live four-major CI lane. The external relay — the default path, and the one most evaluations will actually deploy — has none. Its PostgreSQL integration tests are gated behind TRELLARA_*_DATABASE_URL and the main workflow starts no Postgres service. Deterministic simulations run on every change; version-matrix integration testing of the relay does not.
Support matrix internally consistent
noThree numbers that should agree do not. The runtime release contract advertises external PostgreSQL 16–18 and native 17–18; extension CI builds, live-tests and packages 15–18; the compose harness runs 16. We publish all three rather than the flattering one. Reconciling them is a release gate, not a copy edit.
Managed-provider evidence
noneRDS, Aurora, Cloud SQL, Azure Database for PostgreSQL and Neon differ in extension availability, replication grants, slot behaviour on failover and what they let you observe. Nothing here has been run against any of them. The read-only diagnostic uses ordinary TLS connection handling so they take the same code path, which is a design choice, not evidence.
Public qualification report
not publishedThe correctness report is generated from the simulations and published, which is the machine-checkable half. A qualification report — the thing a review actually asks for, with a named workload, a duration and an environment behind it — does not exist, because none of those have been run.
Soak run
noneNo multi-day continuous run has been performed. Every correctness result on this page comes from bounded deterministic simulation.
Maximum qualified source count
not establishedThe 2,000-database figures on this page are an illustration of the design target. The largest fleet actually exercised is a simulation, not real databases.
Recovery-time distribution
not measuredRecovery is proven correct in 63 scenarios. It has never been timed. We can tell you it converges; we cannot yet tell you how long it takes.
WAL retention under real outage
modelled onlySlot pressure and retention behaviour are modelled in simulation and graded by the diagnostic. Neither has been observed against a production source under a real multi-hour outage.
Upgrade and rollback evidence
noneThere is no released version to upgrade from, so there is no upgrade or downgrade path to have tested.
Signing, SBOM, provenance
checksums onlyRelease artifacts carry SHA-256 checksums. They are not signed, there is no SBOM, and there is no build-provenance attestation.
External correctness review
noneNo third-party analysis of any kind. The correctness argument is entirely our own, which is exactly why it is published as runnable scenarios rather than as prose.
Named design partners
none yetNo production deployments and no reference customers. When there are, they will be named here rather than shown as an unattributed logo.
Throughput and latency
deliberately unmeasuredNot a gap we intend to fill with a synthetic benchmark. Numbers will appear here after a live workload run, and not before.
Compliance and support
noneNo SOC 2, no managed cloud, no support SLA, no non-PostgreSQL sources. If you need a vendor with a signed uptime commitment today, that is a real reason to choose someone else.

# nothing above is a projection · every left-hand row is countable in the repository or produced by CI today
# every right-hand row will move left only when the thing itself exists — not when the copy improves

04 · Reference architecture

Where it runs, what it touches, and where the pressure goes.

Trellara runs entirely inside your network as ordinary processes. On the default path it speaks the standard PostgreSQL replication protocol, so there is nothing to install in the database and nothing to change about the Postgres you already operate. An optional native extension exists for self-managed servers that permit a preloaded library; it is a separate decision with a separate risk profile, and nothing below depends on it.

Your PostgreSQL — unmodified

source · store_001

one logical slot, pgoutput

source · store_002

one logical slot, pgoutput

… source · store_N

RDS · Aurora · Cloud SQL · Azure · self-hosted

START_REPLICATION over the standard protocol · no extension, no superuser-owned agent, no data at rest outside your network

trellara relay Shipping

Decodes pgoutput protocol v2 with streaming, assembles committed transactions into a versioned envelope — boundaries, commit LSN, schema fingerprint, row sequence, checksum — and advances the source checkpoint only after the publish is durable.

publish durably, then acknowledge the source — never the other way round

local durable log Shipping

Disk-backed segment log with fsync, offset index, per-consumer cursors, torn-tail detection and large-transaction spill. This is where retention pressure lives instead of your primary.

Kafka / Redpanda Shipping

The scale-out transport when you already run a broker. Production contract requires TLS, RF≥3, acks=all, idempotence, and offsets stored only after the target apply commits.

one capture, many materializations

Postgres applier Shipping

Applies each source transaction atomically. Data and checkpoint commit together, so redelivery is detected rather than re-applied.

Iceberg fan-in writer In progress

Append-only raw changelog to an Iceberg REST catalog on S3-compatible storage, epoch-batched with per-file digests. Current-state and SCD2 stay Spark templates by design.

your consumers Shipping

The envelope is an open protocol carried as a protobuf IDL in the repository. Search indexes, caches and custom sinks read the same verified stream.

every destination reports a watermark

proof plane Shipping

Source and target watermarks, primary-key checksums, snapshot-handoff state, quarantine records, fleet completeness epochs, and an exportable evidence package. Reached through the CLI, Prometheus metrics, and JSON for your own automation.

Deliberately absent: no transformation engine, no scheduler, no catalog of its own, no phone-home telemetry, and no requirement to adopt a different PostgreSQL distribution. Trellara moves what Postgres committed and proves where it landed. Everything else is somebody else's product.

Failure and recovery: relay dies after publishing, before acknowledging the source

The hardest of the common failures, because it is the one where a naive implementation loses data. Trellara chooses the other failure mode.

Sequence diagram: relay crash after publish, before source acknowledgement The relay reads changes from PostgreSQL, assembles a committed transaction into an envelope and publishes it durably to the log. It then crashes before acknowledging the source. On restart it resumes from the last durable source checkpoint and re-publishes, producing a duplicate by design. The applier applies the transaction atomically, detects the redelivery by idempotency key so the effect happens exactly once, and verification confirms matching watermarks and checksums. PostgreSQL relay durable log applier target START_REPLICATION · pgoutput v2 assemble committed transaction → envelope commit LSN · schema fingerprint · row sequence · checksum publish · fsync relay dies here the publish is durable · the source has not been acknowledged restart · resume from last durable checkpoint re-publish duplicate in the log — by design deliver redelivery detected by idempotency key at-least-once delivery · exactly-once effect apply transaction atomically data and checkpoint commit together verify · watermarks match · primary-key checksums match CONVERGED

The choice that makes this work is ordering: the publish is made durable before the source checkpoint advances. A crash in the window between them costs you a duplicate, never a gap — and the applier is built to absorb duplicates transactionally. Reversing that order would make the happy path marginally cheaper and would lose committed transactions on exactly this failure.

2 of 63 deterministic scenarios cover this window trellara chaos run recovery: automatic on restart

05 · Proof

We rehearse the crashes you are afraid of. Deterministically, on every change.

Every claim above maps to scenarios in a deterministic failure matrix that runs in CI and ships in the binary, so you can rehearse the same crashes against your own stack before production does it for you. Each scenario names its protected invariant, its proof command, and its recovery path.

pass

Relay killed before publish

Nothing acknowledged, nothing lost. Capture resumes from the durable checkpoint.

pass

Relay killed after publish

Duplicates on replay, by design. Zero double-applies at the target.

pass

Applier killed mid-transaction

The target transaction rolls back whole. Data and checkpoint commit together.

pass

Applier killed after commit

Redelivery is detected by idempotency key. The effect happens exactly once.

pass

Broker restart mid-stream

Ambiguous acknowledgements resolved safely; source feedback waits for durability.

pass

Duplicate replay storm

At-least-once delivery, exactly-once effect. Deduplication is transactional.

pass

Partition barrier incomplete

Cross-partition transactions are held, never half-exposed to a reader.

pass

Pinned schema drift

A changed relation fingerprint stops capture before rows are assembled under an unexpected schema.

$ 63 scenarios · 28 deterministic simulations · trellara chaos run · trellara chaos report generates the full HTML correctness report  —  run it against your own stack

Every scenario here is deterministic simulation. That is a real and unusually strong form of evidence, and it is not the same as a soak run against production hardware — which we have not done. The full list of what we have not established →

You are going to ask how this compares. So we wrote it down.

An honest, source-cited comparison against native PostgreSQL logical replication through PG 19, against Striim, and against the CDC and ingestion tools you are probably already evaluating — including a plain statement of where each of them is the better answer, and a list of what Trellara does not have.

Native logical replication Striim Debezium + Kafka Connect Fivetran · Airbyte · Estuary pgEdge Databricks Lakebase

06 · Security & deployment posture

Written for the security review, not around it.

The questions an architecture review actually asks, answered before you have to ask them.

Where it runs

Ordinary processes in your VPC, your Kubernetes cluster, a store back-office server, or a laptop. Distributed as a tiny check binary, a default CLI, a full CLI, and a container image.

no control plane to call home to · no managed dependency

What it installs in your database

On the default path, nothing. Trellara uses the standard replication protocol with pgoutput, which is the only decoding plugin guaranteed present on managed PostgreSQL. A publication and a logical slot are ordinary database objects you can inspect and drop.

no extension · no superuser agent · nothing preloaded

The optional native extension In progress

For self-managed servers that permit shared_preload_libraries, an opt-in extension moves decoding in-process. Broker and network I/O still stay outside PostgreSQL. It is a separate decision with a separate risk profile — you never need it to evaluate or to run the default path.

Deb + RPM for PG 15–18 · bounded 16-slot queue · no two-phase commit

What the first command can do

trellara-check is strictly read-only. It creates no publication, no slot, no checkpoint and no target state, and it is safe to run against production on an evaluation call.

inspection only · refuses to write config when critical blockers exist

Where your data goes

Between your source, your durable log or broker, and your destinations. Nothing is sent to us. There is no vendor telemetry and no hosted component in the open-source path.

object storage you own · catalog you own · engine you choose

Secrets and transport

Credentials are referenced by environment-variable name or file path — never inline in config. The Kafka production contract requires TLS with mutual TLS or SASL, RF≥3, and min.insync.replicas≥2.

no inline secret fields · contract-enforced, not documented-and-hoped

What you can observe

Prometheus metrics with contract-locked names and exactly three labels — service, source, dataset. Transaction ids, LSNs and error strings are forbidden as labels; they go to structured logs and evidence records.

no unbounded cardinality by design

What the evidence artifacts contain

Reports, proof bundles and pilot packages are designed to be forwardable: no passwords, tokens, private keys, certificate contents or complete connection strings. Verification output can contain bounded row samples by construction — that is what makes it evidence — so treat a package as data-classified at the level of the tables it covers.

redaction tested per output surface · retention policy is yours to set

Supply chain

Release binaries are built in CI from a pinned Rust toolchain and published with SHA-256 checksums. They are not signed, there is no SBOM, and there is no provenance attestation yet.

checksums today · signing, SBOM and provenance are open gaps

Licensing and lock-in

Apache-2.0. The transaction envelope is a protobuf IDL carried in the repository rather than a private encoding, the durable log format is documented frame by frame, and the Iceberg output is standard Parquet under your catalog. If you stop using Trellara, your data stays readable and anything you built against the envelope keeps working.

every independent Postgres CDC company so far was absorbed — neutrality is the design response

What we do not offer yet Honest

No managed cloud, no support SLA, no SOC 2, no third-party correctness audit, and no non-PostgreSQL sources. Pre-1.0 with a public roadmap. If you need a vendor with a signed uptime commitment today, that is a real reason to choose someone else.

the itemised list is in §03 · the competitive version is on the comparison page

07 · Evaluation path

Four steps, each one producing an artifact you can take to a review.

No demo environment, no sandbox data. Every step runs against your own PostgreSQL and ends in something you can forward to someone who was not on the call.

60s

Assess a real source

Read-only. Grades slot posture, WAL retention risk, capture readiness and failover-slot guidance on a database you already run.

trellara-check $POSTGRES_URL --format text

→ a graded finding about your system

15m

A verified flow on a laptop

Brokerless local config, snapshot, bounded relay and apply, then a convergence check — the whole loop with no infrastructure to stand up. This is the documented public loop, in order.

trellara init --source-database-url $POSTGRES_URL --table public.sales --evaluate
trellara check --config trellara.yml --format text
trellara preflight --config trellara.yml
trellara run --config trellara.yml --local --verify --format text
trellara verify --config trellara.yml
trellara status --config trellara.yml --view report --format text

→ CONVERGED, or a named recovery path

1 day

Break it on purpose

Run the failure matrix against your own configuration and generate the correctness report. This is the step that replaces trust with evidence.

trellara chaos run
trellara chaos report --output report.html

→ 63 scenarios, pass/fail, per-invariant

2 wks

Pilot with an evidence package

A scoped pilot with a scorecard, recovery drills and an exportable package: convergence evidence, contract tests, drill results, failure matrix.

trellara pilot-scorecard --config trellara.yml --format text
trellara pilot-package --config trellara.yml --output ./pkg

→ the artifact procurement actually wanted

Worth knowing before you type them: trellara --help shows nine commands — init, check, preflight, run, verify, status, fleet, lake and config. Step 2 is exactly that public loop. The chaos and pilot-* commands in steps 3 and 4 work, and are tested, but are hidden from the short help surface as operator and evidence tooling rather than supported public CLI. We would rather tell you that than have you find out.

08 · Who runs it

For teams where PostgreSQL is the business, and there is more than one of it.

vertical saas & isvs

A database per customer

Clinic software, restaurant platforms, field service, practice management. Hundreds of customer databases become one operated fleet with per-tenant replay, reseed and verification.

“When a customer's security review asks you to prove tenant isolation, what do you actually show them?”

retail & logistics

Stores that go dark

Point-of-sale and warehouse databases disconnect for hours at a time. Sales flow when they return, in order — and tonight's numbers carry a completeness statement rather than an assumption.

“At 6 a.m., which stores are in the revenue number and which are missing?”

ai & agent platforms

Ten thousand tenant databases

A platform running a Postgres per tenant is structurally a fleet. An agent that reads a half-applied transaction and then issues a refund is worse off than one reading data two minutes old.

“Can you prove the tenant boundary held while you built the fleet-wide view?”

platform & data teams

One stream, every consumer

Replicas, search, caches and Iceberg tables served from a single verified transaction stream — one slot on each source, watermarks everywhere else, one place to look when something diverges.

“How many slots are open on that primary right now, and who owns each one?”

09 · Where this lands

The wedge is verified replication. The destination is a completeness layer for Postgres fleets.

Published so you can judge the direction, and so you can hold us to it. Items move left only when the command exists and the failure matrix covers it.

Now

Shipping
  • Source safety as the entry point — read-only, graded, actionable
  • Postgres→Postgres verified replication with atomic transaction apply
  • Brokerless durable log so the first run needs no infrastructure
  • Convergence verification and the exportable evidence package
  • Schema contracts, DDL propagation barriers and quarantine
  • 63-scenario failure matrix in CI and in the binary
  • Optional native extension, live-tested and packaged for PG 15–18 In progress

Next

In progress
  • A tagged v0.1.0 with signed artifacts, an SBOM and build provenance
  • A published qualification report at a stable URL, rebuilt from CI
  • A PostgreSQL version matrix exercised in CI, not by hand
  • Fleet runtime — many sources operated as one object, not many configs
  • Iceberg fan-in in production — append-only changelog, REST catalog, S3
  • A CI Postgres lane for the external relay — the native extension already has one across four majors; the default path does not
  • One reconciled support matrix — the runtime contract, the extension CI matrix and the compose harness currently disagree
  • Design partners — the next milestones get chosen by evidence, not inference

Then

Designed
  • Tenant-boundary attestation for platforms with a database per customer
  • Audit-grade provenance — append-only lineage against the 2027 EU AI Act logging duties
  • Fleet-wide schema rollout orchestration with per-sink acknowledgement
  • Managed control plane — only once operators have told us what it should contain
  • Point-in-time views derived from the same verified history

Design partner program — early access

Your database keeps its promises. Your replication layer should be able to prove it kept them too.

We are onboarding a small group of teams with real PostgreSQL fleets and real replication scars. This is a pre-1.0 open-source project with an unusually strong correctness story and an unusually honest roadmap. If that combination is interesting, the first step costs you sixty seconds and a read-only connection string.

What the program actually is

You bring
A real fleet — more than one PostgreSQL database, ideally with nodes that go offline — and an hour to answer five questions about how you operate it today.
We bring
A read-only assessment of your sources, a scoped pilot, direct access to the people writing the code, and roadmap influence that is real because there are not many of you yet.
Cost
Nothing. The software is Apache-2.0 and there is no commercial offering to sell you.
The catch
It is early. There is no SLA, no managed cloud and no support contract. You would be evaluating a correctness argument, not buying a finished platform.