6fbea4fe6f7dcd7ecd8dafea7ba16e88f542f29e
9
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2b695d6892 |
fix(appview): Unfollows über den Firehose anwendbar machen
Ein Delete-Event trägt nur did + rkey, keinen Record-Body. `follows` hatte aber nur (follower_did, subject_did) und speicherte den rkey nicht — es gab also keinen Weg vom rkey zum subject_did, und der Indexer hat solche Ops geloggt und übersprungen. Unfollows hingen damit allein am Best-Effort-Push, genau der Abhängigkeit, die der Firehose beseitigen soll. Migration 0011 ergänzt die rkey-Spalte plus einen partiellen Index für den Lookup. Der Primärschlüssel bleibt (follower_did, subject_did), damit die Upserts über Push, Firehose und Replay hinweg idempotent bleiben; ein rkey im Schlüssel würde aus einem Re-Follow eine zweite Zeile machen und die Follower-Zahl verdoppeln. Der Index ist bewusst nicht unique: sonst würde ausgerechnet der Fall, für den das hier existiert — verlorener Delete, dann ein neuer Create — zu einem abgebrochenen Write. delete_follow_by_rkey löst und löscht in einem Statement (RETURNING), also ohne Rennen zwischen Auflösen und Löschen. Findet es nichts — alte Zeile ohne rkey, schon gelöscht, veralteter rkey — ist das kein Fehler. Der Push-Pfad über subject_did bleibt unverändert. Likes und Reposts haben die Lücke nicht: dort ist der rkey Teil der Zeilenidentität. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX |
||
|
|
124a90dc07 |
feat(appview): PDS-Firehose konsumieren
Gegenstück zum subscribeRepos-Endpoint: WebSocket-Consumer mit persistiertem seq-Cursor, Reconnect-Backoff und Behandlung von #info/OutdatedCursor. Eigene Cursor-Tabelle statt einer Zeile in jetstream_cursor: dort steht ein time_us in der Größenordnung 1.7e15, die seq ist ein kleiner Zähler ab 1. Geteilt hätte GREATEST den PDS-Cursor sofort in eine Zukunft geschoben, die die PDS nie erreicht. Kein neuer Indexer-Pfad — jede Op wird in die Single-Op-Form übersetzt, die apply_commit schon vom Jetstream kennt. Push und Firehose liefern denselben Commit doppelt; das ist unkritisch, weil die Schreibpfade Upserts sind und der Dedupe-Index der Notifications den Rest abfängt. Mit einem Test festgehalten statt vorausgesetzt. Der CAR-Reader ist neu (es gab nur einen Writer, und der liegt in einem Binary-Crate ohne lib-Target). Der CBOR-Reader arbeitet mit explizitem Offset, weil ein Frame zwei hintereinander geschriebene Werte sind, und akzeptiert CID-Links in beiden Schreibweisen — die Blöcke tragen Strings. /healthz meldet beide Ströme getrennt; sie fallen unabhängig voneinander aus. Verifiziert mit totem Push-Ziel: der Post kam trotzdem an. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX |
||
|
|
0646fbeebe |
feat(pds): com.atproto.sync.subscribeRepos — lokaler Firehose
Bisher erreichten eigene Records die AppView nur über den Best-Effort-Push /internal/ingest-commit. Ging der verloren (AppView kurz weg, Netzwerk- fehler), war der Post dauerhaft weg: der öffentliche Jetstream kennt diese PDS nicht, es gab also keinen zweiten Weg. Jeder Commit schreibt sein Event in derselben Transaktion nach firehose_events. Damit kann es keinen Commit ohne Event geben — und keine Sequenz ohne Commit. Die seq muss lückenfrei sein, sonst ist sie als Cursor wertlos: BIGSERIAL vergibt Nummern bei INSERT, nicht bei COMMIT, also können zwei Schreiber 5 und 6 ziehen und in umgekehrter Reihenfolge sichtbar werden — ein Leser dazwischen sieht 6, merkt sich das und erfährt von 5 nie. Ein globaler pg_advisory_xact_lock unmittelbar vor dem INSERT erzwingt Commit-Reihenfolge == seq-Reihenfolge. Er wird nach dem per-Repo-FOR-UPDATE genommen, überall in derselben Reihenfolge, also ohne Deadlock-Risiko. Preis: das Ende jeder schreibenden Transaktion ist global serialisiert; das steht im Modulkopf. Der WebSocket-Handler abonniert den Broadcast, *bevor* er die Datenbank liest, und filtert Live-Events auf seq > Wasserstand. Aus einem Rennen wird so eine Dublette, die sich filtern lässt, statt einer Lücke, die es nicht gibt. Ein zu langsamer Consumer bekommt #info/OutdatedCursor und fällt auf den DB-Replay zurück, statt getrennt zu werden — die Events sind durabel, also ist der Rückfall verlustfrei. Frame-Hülle ist konformes DAG-CBOR mit Tag-42-Links (neues Modul dag_cbor, aus car.rs herausgezogen statt dupliziert). Die Blöcke darin behalten die Konvention dieses Repos: CIDs als Strings. Ein fremder Consumer liest die Frames, scheitert aber an den Blockinhalten — das zu ändern hieße, jede CID im System zu ändern, inklusive der did:plc-Ableitung. Steht so im Modulkopf. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX |
||
|
|
73da8f0140 |
perf(appview): Profil, Cold-Start-Feed und Follow-Timeline entlasten
Gemessen gegen die Dev-Instanz (3,3 Mio. Posts): * GET /api/profile/<handle> 9,5 s → 0,04 s * Cold-Start-Timeline 7,4 s → 0,006 s * Timeline mit 2300 Follows 28 s → 0,02 s Drei unabhängige Ursachen, alle drei ein Seq-Scan über die posts-Tabelle: 1. resolve_profile sucht die DID über profiles.LOWER(handle) und, als Fallback, über posts.handle. Für beides gab es keinen Index. Auf profiles hatte Migration 0007 genau diesen Index entfernt, mit der Begründung, jeder Aufrufer leite ohnehin zuerst eine DID ab — das stimmt nicht mehr, seit resolve_profile den profiles-Cache zuerst befragt. 2. Der Cold-Start-Feed filtert `collection IN (…)` und sortiert nach indexed_at. Der vorhandene (collection, indexed_at, uri)-Index taugt dafür nicht: mit zwei führenden Werten liefert er keine indexed_at-Ordnung mehr. Ein partieller Index über genau das Prädikat schiebt den Filter in die Definition und lässt (indexed_at DESC, uri DESC) als Sortierschlüssel übrig. 3. Genau dieser neue Index wurde dann zur Falle für den Graph-Zweig: der Planer sah einen Index, der schon in indexed_at-Ordnung liefert, und nahm an, er treffe früh genug auf n passende Zeilen — bei dünn besetzten Followees hieß "früh" 2,87 Mio. verworfene Zeilen. Je nach Anzahl bisheriger Ausführungen des Prepared Statements kippte er zwischen diesem und dem guten Plan, was intermittierend aussah. Der Graph-Zweig formuliert die Absicht jetzt aus: pro Followee die neuesten Posts über ein LATERAL, dann mergen. Damit ist der globale Scan kein wählbarer Plan mehr, und jede Iteration ist ein begrenzter Range-Scan auf posts_did_indexed_at_uri_idx. Korrekt ist das, weil die globalen Top-N immer eine Teilmenge der Vereinigung der Top-N je Followee sind — deshalb wird pro Followee limit+1 geholt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX |
||
|
|
c4ca218d97 |
feat(appview): Notifications, Follower-/Following-Listen, Thread-Route
Bisher erfuhr ein Nutzer nie, dass jemand anderes mit ihm interagiert hat:
Like, Repost, Follow und Reply hinterließen keine Spur, an der der Client
hätte pollen können. Der Tray-/Notification-Pfad im Desktop-Client (Phase 7)
hing damit in der Luft.
Migration 0008:
* notifications(recipient, author, kind, subject_uri, created_at,
indexed_at, read_at) mit Keyset-Index (recipient, indexed_at DESC, id DESC)
und Partial-Index auf ungelesene Zeilen für den Badge-Poll.
* Dedupe-Unique-Index über COALESCE(subject_uri, '') — plain NULLs
kollidieren nicht, sonst gäbe es pro Follow beliebig viele Zeilen.
Folge: Unlike-Relike erzeugt keine zweite Notification, Toggle-Spam ist
damit ausgeschlossen.
* Bewusst kein CHECK (recipient <> author): ein Ausrutscher dort würde die
umgebende Like-Transaktion abbrechen, also das Like wegen eines
Notification-Bugs verlieren. Gefiltert wird in Rust und im INSERT.
Indexer: record_notification() hängt an upsert_like/-repost (in derselben
Transaktion wie die Counter) sowie upsert_follow/-post. Selbst-Interaktionen
sind still. Empfänger muss uns bekannt sein (profiles- oder posts-Zeile),
sonst würden wir für den gesamten öffentlichen Firehose Zeilen anlegen —
als ein INSERT ... SELECT ... WHERE EXISTS, also ohne TOCTOU-Fenster.
Reply-Notifications tragen die URI der *Antwort* als subject_uri, weil die
Liste den Text zeigt, den der Empfänger noch nicht kennt.
Endpoints: GET /api/notifications, /api/notifications/count,
POST /api/notifications/seen (seenAt als Wasserzeichen),
GET /api/followers, /api/following, GET /api/thread (beide Schreibweisen).
Cursor-Codec, Limit-Clamping und Fehlerform sind die der bestehenden
Endpoints.
/api/post/*uri bleibt wire-kompatibel und teilt sich jetzt
load_thread_context() mit /api/thread — mit max_parents = 1, weil es nur
den direkten Parent serialisiert; die volle Ahnenkette wären bis zu 20
sequenzielle Queries für Zeilen, die danach verworfen werden.
Nebenbei ein Darstellungsfehler: der synthetische Platzhalter-Handle für
Actors ohne bekannten Handle trug ein führendes '@', während jeder Consumer
selbst '@{handle}' rendert — im Feed kam '@@did:plc:abcd…' heraus. Der
Platzhalter ist jetzt durchgängig sigil-frei.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013HC9HLrUU1LNwkzp8nkDLX
|
||
|
|
6ebf17b493 |
fix(appview): resolve_profile also looks up DID in profiles cache
A user who set their profile via the PDS push path before posting anything has a row in `profiles` but no rows in `posts` — the handle→DID lookup in `resolve_profile` only checked `posts`, so the route synthesised an empty profile (`did: ""`) for these users. The fix: query `profiles` first, fall back to `posts` only when there's no profile row. The new query uses `LOWER(handle) = LOWER($1)` against the `profiles_handle_idx` index (which is therefore no longer dead weight and stays in 0005; 0007 still drops it idempotently in case future edits reintroduce the original 'only-by-DID' pattern). Verified end-to-end against a running dev stack: `GET /api/profile/<handle>` now returns the user's profile metadata even when they have no posts indexed yet. |
||
|
|
59a3cb02dd |
feat(appview): profile cache + Jetstream indexing + denormalised counts
Adds the AppView-side half of the profile feature so non-local-PDS
authors also get their profile metadata indexed (the Jetstream
identity event stream only carries the handle, not display name /
bio / avatar). The PDS-push path was already wired by the previous
commit; this lands the Jetstream path.
Migration 0005:
* `profiles` table keyed by DID with display_name / description /
avatar_cid / banner_cid plus denormalised post_count /
follower_count / following_count. Backfilled from the posts
table on apply.
* `posts.avatar_cid` column — populated from the profiles cache
at `upsert_post` time so the PostCard can render an avatar
inline without a per-row PDS round trip.
Migration 0007 (clean-up): the original 0005 also created a
`LOWER(handle)` index that no query uses; this drops it
idempotently so dev DBs that already applied 0005 converge.
Indexer (`crates/appview/src/indexer.rs`):
* New `app.bsky.actor.profile` arm in `apply_commit` calls
`upsert_profile` on create, DELETEs the row on delete. Handle
is looked up from `posts` (the Jetstream commit envelope
doesn't carry it).
* `upsert_post` signature is now `&mut PostRow` so it can fill
`row.avatar_cid` from the profiles cache; the ON CONFLICT
clause uses `COALESCE(EXCLUDED, posts)` so re-indexing doesn't
overwrite an already-known avatar.
* `upsert_profile` writes display_name / description /
avatar_cid / banner_cid + the denormalised counts.
* `blob_link_of` helper accepts both `{ $type, ref.$link }`
and legacy flat `{ $link }` blob-ref shapes.
Ingest (`crates/appview/src/ingest.rs`):
* `app.bsky.actor.profile` create/delete arms in the PDS-push
path. The handle-fallback previously did `SELECT handle FROM
users WHERE did = $1` — but the AppView has no `users` table
(it's PDS-owned state). Replaced with a simple use-what-the-PDS-
sent approach; the handle_sync worker fills the column later.
Routes (`crates/appview/src/routes.rs`):
* `resolve_profile` reads the denormalised profile fields from
the cache. When no profile row exists the `post_count` fallback
uses a live `SELECT COUNT(*)` instead of `posts.len()`, so
prolific authors without a profile row report the real count
rather than the 50-post slice cap.
Tests (DB-gated, run when DATABASE_URL_APPVIEW is set):
* `blob_link_of_modern_shape` / `_legacy_flat_link` /
`_missing_field`.
* `upsert_profile_round_trip` — insert + replace semantics.
* `apply_commit_indexes_profile_create` — end-to-end Jetstream
arm + delete.
|
||
|
|
3aa5d5c0e3 |
feat(at-identity): PdsHandleResolver for cluster-local DID→handle resolution
The AppView's handle_sync worker consulted the public PLC directory and the did:web: HTTPS resolver only. DIDs hosted on the local PDS (notably did🔑 users and any other operator-hosted method) weren't reachable without an external round trip, and unresolvable DIDs (did🔑 not on this PDS, did:foo: anything) blocked the 100-row batch forever because did🔑 sorts lexicographically before did:plc: / did:web:. This commit adds: * `PdsHandleResolver` (at-identity) — POSTs the DID as the `handle` field to the PDS's resolveHandle XRPC method. The PDS now recognises a `did:` prefix and does a PK lookup on `users.did`, returning `{did, handle}`. The resolver reads the `handle` field, so the AppView finally gets a real local handle for did🔑 users without ever dialing plc.directory. * A 2 s timeout per request (was 10 s) and `DISPATCH_CONCURRENCY = 8` so the worker caps a 100-DID batch at ~2 s with parallel dispatch instead of the ~17 min worst case the old serial + 10 s setup allowed. * A new `posts.handle_sync_attempted_at` column (migration 0006) and `mark_attempted()` helper. The SELECT filter excludes rows attempted within the last hour, so an unresolvable DID dominates at most one batch before the worker advances. Cleared on success. * `PDS_INTERNAL_URL` config so the AppView can reach the PDS via a cluster-internal hostname when the public URL isn't routable from inside the cluster. Tests: * `crates/at-identity/src/pds_handle.rs` — 4 unit tests against a stub HTTP server (200/404/5xx/missing-did-field). * Existing handle_sync integration tests updated to wire in the new `pds_resolver` field. |
||
|
|
c586fd39c9 |
maarcadetweet: initial commit
AT Protocol PDS + AppView + Tauri Desktop Client, 160-char post limit. - PDS (Rust + axum + sqlx) - Auth: createAccount, createSession, refreshSession - Records: createRecord, deleteRecord (race-safe via SELECT FOR UPDATE) - Feed: feed.like.create, feed.repost.create - Sync: getRepo, getBlocks, getLatestCommit, getRecord (with MST proof), listRepos - Identity: resolveHandle - MST: spec-conformant (at-mst crate, 27 tests) - Repo: signed commits, TID counter (monotonic, 4096 wrap safe) - AppView (Rust + axum + sqlx) - Jetstream consumer (WebSocket, exponential backoff, 38k+ events indexed) - REST API: timeline/home (graph-aware), profile, search, post (with thread hydration) - Handle-sync worker (did:plc + did:web) - JSONB embed storage + thread columns (migration 0003) - Like/repost counter cache (migration 0004) - Tauri 2 + Svelte 5 Desktop Client - System tray (Show/Compose/Quit menu) - OS notifications (tauri-plugin-notification) - Auto-update (tauri-plugin-updater, placeholder endpoint) - Window-state (tauri-plugin-window-state) - 160-char compose with live counter - Image/Link embed rendering - LocalStorage-persisted like state - Timeline with poll (prepend new posts) - Custom TitleBar (transparent, no decorations) - Orange/IBM Plex Mono maarcade design Tests: 231 Rust + 9 vitest = 240 passed. |