feat(at-identity): PdsHandleResolver for cluster-local DID→handle resolution

The AppView's handle_sync worker consulted the public PLC directory
and the did:web: HTTPS resolver only. DIDs hosted on the local PDS
(notably did🔑 users and any other operator-hosted method)
weren't reachable without an external round trip, and unresolvable
DIDs (did🔑 not on this PDS, did:foo: anything) blocked the
100-row batch forever because did🔑 sorts lexicographically
before did:plc: / did:web:.

This commit adds:

* `PdsHandleResolver` (at-identity) — POSTs the DID as the
  `handle` field to the PDS's resolveHandle XRPC method. The PDS
  now recognises a `did:` prefix and does a PK lookup on
  `users.did`, returning `{did, handle}`. The resolver reads
  the `handle` field, so the AppView finally gets a real local
  handle for did🔑 users without ever dialing plc.directory.
* A 2 s timeout per request (was 10 s) and `DISPATCH_CONCURRENCY =
  8` so the worker caps a 100-DID batch at ~2 s with parallel
  dispatch instead of the ~17 min worst case the old serial + 10 s
  setup allowed.
* A new `posts.handle_sync_attempted_at` column (migration 0006)
  and `mark_attempted()` helper. The SELECT filter excludes rows
  attempted within the last hour, so an unresolvable DID dominates
  at most one batch before the worker advances. Cleared on success.
* `PDS_INTERNAL_URL` config so the AppView can reach the PDS via
  a cluster-internal hostname when the public URL isn't routable
  from inside the cluster.

Tests:
* `crates/at-identity/src/pds_handle.rs` — 4 unit tests against a
  stub HTTP server (200/404/5xx/missing-did-field).
* Existing handle_sync integration tests updated to wire in the
  new `pds_resolver` field.
This commit is contained in:
tomdebone
2026-07-18 17:56:11 +02:00
parent 391448a845
commit 3aa5d5c0e3
11 changed files with 378 additions and 27 deletions
@@ -0,0 +1,25 @@
-- AppView database schema 0006: handle-sync attempt tracking.
--
-- Why
-- The `handle_sync` worker SELECTs DIDs whose `posts.handle` is empty
-- and tries to resolve them via the local PDS → PLC directory →
-- `did:web:` resolver. Some DIDs are *unresolvable* (e.g. a `did:key:`
-- user not hosted on the local PDS, or any `did:foo:` method that
-- neither PLC nor Web understands). Without tracking these, every
-- pass re-selects them and they dominate the 100-row batch — and
-- since `did:key:` sorts lexicographically before `did:plc:` /
-- `did:web:`, the worker would process the same 100 unresolvable
-- `did:key:` rows forever and never reach any resolvable DID.
--
-- With this column the worker marks each empty-handle row with the
-- time of its last attempt. The SELECT filter excludes rows
-- attempted within the last hour, so an unresolvable DID gets at
-- most one attempt per hour and stops blocking forward progress.
-- Rows whose `handle` later gets filled (by another code path) are
-- naturally no longer in the candidate set.
--
-- The column is per-post (not per-DID) because the candidate set is
-- already per-post and the update is cheap (the empty-handle slice
-- is small in steady state).
ALTER TABLE posts ADD COLUMN IF NOT EXISTS handle_sync_attempted_at TIMESTAMPTZ;