feat(pds): com.atproto.sync.subscribeRepos — lokaler Firehose

Bisher erreichten eigene Records die AppView nur über den Best-Effort-Push
/internal/ingest-commit. Ging der verloren (AppView kurz weg, Netzwerk-
fehler), war der Post dauerhaft weg: der öffentliche Jetstream kennt diese
PDS nicht, es gab also keinen zweiten Weg.

Jeder Commit schreibt sein Event in derselben Transaktion nach
firehose_events. Damit kann es keinen Commit ohne Event geben — und keine
Sequenz ohne Commit.

Die seq muss lückenfrei sein, sonst ist sie als Cursor wertlos: BIGSERIAL
vergibt Nummern bei INSERT, nicht bei COMMIT, also können zwei Schreiber 5
und 6 ziehen und in umgekehrter Reihenfolge sichtbar werden — ein Leser
dazwischen sieht 6, merkt sich das und erfährt von 5 nie. Ein globaler
pg_advisory_xact_lock unmittelbar vor dem INSERT erzwingt Commit-Reihenfolge
== seq-Reihenfolge. Er wird nach dem per-Repo-FOR-UPDATE genommen, überall in
derselben Reihenfolge, also ohne Deadlock-Risiko. Preis: das Ende jeder
schreibenden Transaktion ist global serialisiert; das steht im Modulkopf.

Der WebSocket-Handler abonniert den Broadcast, *bevor* er die Datenbank
liest, und filtert Live-Events auf seq > Wasserstand. Aus einem Rennen wird
so eine Dublette, die sich filtern lässt, statt einer Lücke, die es nicht
gibt. Ein zu langsamer Consumer bekommt #info/OutdatedCursor und fällt auf
den DB-Replay zurück, statt getrennt zu werden — die Events sind durabel,
also ist der Rückfall verlustfrei.

Frame-Hülle ist konformes DAG-CBOR mit Tag-42-Links (neues Modul dag_cbor,
aus car.rs herausgezogen statt dupliziert). Die Blöcke darin behalten die
Konvention dieses Repos: CIDs als Strings. Ein fremder Consumer liest die
Frames, scheitert aber an den Blockinhalten — das zu ändern hieße, jede CID
im System zu ändern, inklusive der did:plc-Ableitung. Steht so im Modulkopf.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX
This commit is contained in:
tomdebone
2026-09-10 07:08:02 +02:00
co-authored by Claude Opus 5
parent 6fd046417a
commit 0646fbeebe
14 changed files with 2844 additions and 22 deletions
+112
View File
@@ -0,0 +1,112 @@
-- PDS database schema 0003: the firehose event log.
--
-- Why
--
-- Until now the PDS produced no `com.atproto.sync.subscribeRepos` stream at
-- all. The only way a local record reached the AppView was the best-effort
-- HTTP push in `appview_push.rs` — a fire-and-forget `tokio::spawn` that is
-- explicitly documented as "the Jetstream replay will catch up". There is no
-- Jetstream replay for records that only exist on this PDS, so a dropped push
-- meant the post was simply never indexed. Nothing retried it, and nothing
-- could: the commit lived in `repos` / `repo_blocks` but there was no ordered
-- log of *what changed* for a consumer to walk.
--
-- This table is that log. Every repo write appends exactly one row, in the
-- same transaction as the head-pointer update, so the sequence and the repo
-- head can never disagree. A consumer that reconnects with a cursor replays
-- from here; a consumer that is live gets the same rows pushed over a
-- broadcast channel.
--
-- Column choices
--
-- seq BIGSERIAL PRIMARY KEY — the cursor. It has to be a single
-- monotonically increasing integer because that is what the
-- `subscribeRepos` wire contract hands the client and takes
-- back as `?cursor=`. BIGSERIAL (not an `(timestamp, id)`
-- keyset like the AppView's notifications table) because the
-- protocol's cursor is opaque-but-numeric and clients compare
-- it with `>`.
--
-- Sequence values are handed out at INSERT time, which by
-- itself does NOT guarantee that they become *visible* in seq
-- order — two transactions can grab 5 and 6 and commit in the
-- opposite order, leaving a reader that polls in between with a
-- gap it would never fill. The write path therefore takes
-- `pg_advisory_xact_lock` on a fixed key immediately before
-- this INSERT (see `routes::helpers::apply_repo_write`), which
-- serialises the tail of every firehose-writing transaction so
-- commit order == seq order. That is what makes "give me
-- everything with seq > N" an exact, gap-free replay rather
-- than a best guess.
--
-- did the repo the event belongs to. Not a FK to `users(did)`:
-- the log outlives the account. If a user is deleted we still
-- want consumers that are mid-replay to see the events that
-- already happened rather than have the rows cascade out from
-- under their cursor.
--
-- rev the new commit's revision (TID string), mirrored from
-- `repos.rev`. Goes out as the frame's `rev`.
--
-- since the *previous* commit's rev, or NULL for the first commit on
-- a repo. The frame's `since` field; a consumer uses it to
-- detect that it missed an intermediate commit.
--
-- commit_cid BYTEA holding the raw binary CID of the new commit, stored
-- the same way `repos.head_cid` stores it so the two are
-- directly comparable with `=` and no text/binary conversion
-- is needed to join them.
--
-- blocks BYTEA holding a complete CAR v1 file: the commit block as the
-- root plus every block this commit newly created (MST nodes
-- and record values). Stored pre-serialised rather than
-- reassembled from `repo_blocks` at read time because the
-- *diff* — which blocks were new for this particular commit —
-- is only knowable at write time. Recomputing it later would
-- mean diffing two MST snapshots on every replayed event.
--
-- ops JSONB array of `{action, path, cid}`, the same objects that
-- go into the frame's `ops` field. JSONB rather than a child
-- table because it is always read as a whole, is never queried
-- by content, and a child table would need its own ordering
-- column to reproduce the array faithfully.
--
-- created_at when the event was appended. This is what the frame's `time`
-- field carries, so a replayed frame is byte-identical to the
-- live one that was broadcast at commit time — a consumer that
-- deduplicates by hashing frames does not see two different
-- frames for one event.
--
-- Retention: there is none
-- ------------------------
-- Nothing prunes this table. It grows by one row per repo write, and each row
-- carries a CAR of the commit's new blocks (a few hundred bytes for a plain
-- post, more when a record is large). At the volume this deployment sees that
-- is fine for a long time, but it is unbounded, and an operator who wants a
-- bound has to add one. Deleting the oldest rows is safe: a client whose
-- cursor points before the surviving range gets an `#info`/`OutdatedCursor`
-- frame and resumes from the oldest row that still exists. See the module
-- header of `crates/pds-server/src/firehose.rs`.
CREATE TABLE IF NOT EXISTS firehose_events (
seq BIGSERIAL PRIMARY KEY,
did TEXT NOT NULL,
rev TEXT NOT NULL,
since TEXT,
commit_cid BYTEA NOT NULL,
blocks BYTEA NOT NULL,
ops JSONB NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- Cursor replay is `WHERE seq > $1 ORDER BY seq LIMIT $2`, which the
-- BIGSERIAL primary key's own index already serves — no second index for
-- that, on purpose: an extra index on `seq` would be pure write amplification
-- on the hottest path in this table.
--
-- What the PK does *not* serve is "replay one repo", which is how an operator
-- re-drives a single account into the AppView after an ingest bug, and how
-- `getRepo`-style backfills are debugged. `(did, seq)` covers that and keeps
-- the per-repo scan in seq order.
CREATE INDEX IF NOT EXISTS firehose_events_did_seq_idx
ON firehose_events (did, seq);