Volter World

Conformance

A twin is only worth trusting if its behavior provably tracks the real vendor.

Read this first — conformance answers TWO separate questions

Keep these distinct; conflating them is what produced misleading "100%" claims before.

Completeness Fidelity
Question How much of the real vendor do we cover? Is what we DO model faithful to the vendor?
Measured by the capability manifest vs the vendor's full surface the evidence rungs over the modeled subset
Mechanism <vendor>-capabilities.ts + checkCapabilities (a ground-truth verify() per capability) spec-conformance · recorded-diff · SDK-parity · UI-structure
Output honest coverage % + the ranked, categorized worklist (regression / todo by tier) per-rung pass/fail + deviations over what's built
Run it bun scripts/twin-capabilities.ts bash scripts/twin-check.sh

The trap to avoid: a twin can be high-fidelity but low-completeness — faithful on a small slice. "UI rung ✅" means the mirror faithfully renders what the twin models; it does NOT mean the UI is complete — completeness is the capability %. The honest "how done are we" number is always the capability coverage, never a rung checkmark. Faithfully here means as the product renders it — see What a mirror IS; scope limits what the mirror shows, never how honestly it shows it.

Completeness (the headline + worklist) is described next; the fidelity rungs follow.


Completeness — capability conformance (the headline metric + worklist)

See the dedicated section below ("Capability conformance"). In one line: each world's manifest enumerates the real vendor's full surface; verify() proves each done; coverage is honest and currently partial. This is the number to trust for "how complete is the twin" — see the generated, current per-twin table below (B8: bun scripts/twin-capabilities.ts --write-docs) rather than any percentage quoted in prose, which is a point-in-time snapshot and goes stale the moment a manifest grows.


Fidelity — the ladder of evidence rungs

These rungs prove the quality of what the twin models (not how much it covers — that's completeness, above). From weakest to strongest; each is a different kind of proof, not a replacement for the one below.

# Rung What it proves Oracle
1 Spec-derivation The twin's schema cannot drift from the vendor's, because it's generated from it vendor SDL / OpenAPI (build-time)
2 Spec-conformance Every field the twin emits matches the published spec (type / required / enum), and we know our coverage of the spec published schema
3 Recorded-diff The twin's output matches a real captured response byte-shape (missing / extra / mismatch / type / length) a real vendor response on disk
4 SDK-integration parity The real vendor SDK, unmodified, drives the twin and round-trips the SDK itself
5 UI conformance The UI mirror renders what the real product's UI shows (completeness) and looks structurally like it (proximity) a declared UI surface inventory (+ optional reference screenshot)

Rungs 1–4 cover the API; rung 5 is a parallel dimension covering the UI mirror. A twin states which rungs it holds — see the status table below.

What a mirror IS (read this before building one)

A mirror is the real product's UI, running locally — using it should feel like using the vendor's app, minus the network. It is not a debug view, a data browser, or an inspector over twin state. That is the whole reason it exists: the twin world is what the AI sees and acts on, and the mirror is how a person looks at that world directly and judges whether the agent's version of reality — and everything the agent did to it — is right. A mirror that renders a representation of the data instead of the product breaks that: it looks authoritative while showing something the vendor's UI would never show, so a reviewer signs off on a world they never actually saw.

The concrete rule this implies: for anything the twin models, the mirror shows the real thing, not a stand-in. An image renders as the image, not a labelled empty box. An avatar is the avatar. A file thumbnail resolves. A placeholder is only acceptable where the vendor's own UI shows a placeholder. "The data is present in the DOM as an attribute" is not rendering it — data-image-ref on an empty div is exactly the failure this rule names, because a human reviewing that canvas sees blank rectangles where the design's actual content lives.

This is deliberately stricter than the parity gate below, which only asks that modeled data be shown somewhere. Parity catches omission; this catches misrepresentation, and a mirror can pass parity while failing it.

When serving the real thing needs bytes the twin cannot synthesize (rendered images, uploaded files, generated media), the answer is to serve what the connector pulled — a pulled file carries the vendor's own assets, and the twin's job is to hand them back. Declaring the surface a gap because the twin cannot generate the pixels confuses synthesis with replay: the twin never had to invent them.

A third, orthogonal axis: unit-test depth vs manifest-only coverage. Completeness (above) measures coverage of the vendor's surface; fidelity (this ladder) measures faithfulness of what's modeled. Neither measures how a pack proves its behavior — a hand-authored *.test.ts file vs. the capability manifest's own verify() round-trips (which ARE tests, and are mutation-gated — see scripts/mutation-test.ts in the freshness cycle below). Raw test-file-count is a poor proxy for this (a pack can have few *.test.ts files and a deep, mutation-proven manifest, or vice versa). ./test-depth.md is the per-pack resolution for the eight packs the architecture review flagged as thin by that proxy — each one either gained real test depth or carries a specific, named, written acceptance of manifest-only coverage.

Capability conformance — the exhaustive, categorized gap check

The rung harnesses below measure a twin against its own declarations (a hand-kept inventory/schema). That is self-referential: anything nobody declared is invisible — which is how Linear reported "UI 100%" while missing its entire list view mode. The fix is an authoritative, independent definition of what the real vendor does:

The scope rule: 100% of what the vendor does. There is no "in scope" list and no exclusion list. Everything the real vendor does is covered by default, and every capability is done or todo — nothing else. A twin is a deterministic, offline model of the vendor's API contract; "the real thing" (live model output, real delivery, live data, infrastructure physics, hosted pixels, real external trust) is not a capability a twin lacks, it is the definition of a twin, so it is never listed as a gap. So every gap is a todo — there is no "silently absent" state and no "chose not to" state. This is what stops gaps like Linear's missing list view from hiding — an un-enumerated, not-carved-out capability is, by definition, an open todo.

  • Each world declares a capability manifest (<vendor>-capabilities.ts) — the expected real-product surface: every core capability across API + UI view-modes + connector. This is "what 100% means," reviewed and separate from the implementation.
  • Each capability carries a ground-truth verify() predicate (UI: the built mirror renders the screen; API: the operation runs), so done must be proven, not asserted.
  • Each capability also carries an importance tier — core (the daily-use backbone), common (frequently used), or niche (long-tail / admin / power-user). The tier does not change scope (everything is built or a todo); it only ranks the worklist so we close the highest-value gaps first.
  • checkCapabilities (in @volter/world-tooling — packages/world-tooling/src/capabilities.ts, dev-only, NOT the runtime kernel @volter/world-core) reconciles claim-vs-truth into the exhaustive, categorized, ranked worklist (bun scripts/twin-capabilities.ts):
    • regression — declared done but verify() now fails (a false-green / breakage). Always first.
    • todo — covered by default, not built yet — the worklist, ordered core → common → niche.
    • (done — expected and verify() passes.)
  • A world with no manifest yet is itself the biggest todo (declare it). Each cycle grows every manifest toward the real product's full surface, then closes regression + todo (core todos first). The check produces the ranked list; the closer (subagent) closes it from the top; re-running proves closure. All 69 twins now carry a manifest against their real-vendor surface (honest, verify-proven coverage) — for the current per-twin numbers and total todos, see the generated table below ("Capability coverage") or run bun scripts/twin-capabilities.ts. (Historical snapshot, now stale: at the original 5-world baseline — before this repo grew to 26 twins — coverage read slack 50% (75/150), stripe 46% (58/126), linear 44% (26/59), jira 42% (49/118), github 36% (38/107), 314 todos. Numbers drift every cycle; never treat a prose percentage as current — the generated table is the only one the drift gate keeps honest.)

Adding a twin — conformance checklist (and the anti-patterns that bit us)

A new twin is not done until it satisfies these. The meta-test (scripts/capability-manifests.test.ts, run by twin-check) enforces items 1, 4, 5, 6 mechanically; items 2–3 are judgment and are guarded by adversarial review each cycle.

  1. Declare <vendor>-capabilities.ts — a CapabilitySpec[] + <vendor>Capabilities() calling checkCapabilities. Wire it into scripts/capability-manifests.ts.
  2. The manifest is the REAL vendor's full surface (the target), authored top-down — NOT a list of what you built. A manifest that mirrors your implementation makes the metric meaningless (a thin list inflates coverage). Most entries start todo; coverage should read honestly LOW. (Anti-pattern that bit us: the first manifests were self-portraits → "75%/100%"; rewritten to the real surface → an honest ~30–40%.)
  3. done requires a ground-truth verify() that is a REAL round-trip — create/mutate via the twin → read it back → assert the returned values; negative cases assert the vendor-shaped error. It MUST be failable (would fail if the capability broke). Never verify with "no errors on an empty workspace" / "an empty list parses" / a bundle string present for unrelated reasons. (Anti-pattern that bit us: Linear's dones passed against an empty workspace and proved nothing; Stripe's webhooks.emit only checked a create succeeded. Both were false-greens caught only by adversarial review.) Reference the strong patterns in stripe-capabilities.ts and linear-capabilities.ts (withRoot + create-read-assert).
  4. Enumerate UI VIEW-MODES and screens, not just data fields — list/board/timeline/ calendar, detail panes, inbox, settings, etc. A feature is done only when it's in the API and rendered in the UI mirror (parity). (Anti-pattern: Linear was "UI 100%" while missing its entire list view — nobody had enumerated it.)
  5. Everything the vendor does, ranked by tier. Categories are regression / todo / done. There is no "in scope" list and no exclusion list — scope is 100% of what the vendor does — every gap is a todo, never silently absent and never "chosen away". Tag every capability with a tier (core/common/niche) so the worklist ranks — build core gaps first. Don't report "deep ✅ / 100%" — report the honest measured % against the real surface.
  6. Add <vendor>-capabilities.test.ts asserting assertManifestBaseline(report) (0 regressions, done>0, total>=50).
  7. Assert what DISTINGUISHES the behaviour, not merely that it failed. A verify() that checks only 400 validation_error proves almost nothing on an endpoint with more than one way to reject: delete the guard under test and a different check further down returns the same status, so the capability stays green while the behaviour is gone. Pin the vendor's message (isErrSaying) whenever several rejections share a code. The same trap has a positive form — a fixture where two values coincide. (Anti-patterns, all caught by sabotage-probing rather than by review: Notion's read-only-property, title-removal and version-mismatch guards could each be deleted with every test still passing, because a second path produced the same 400; and its "a content webhook names the PAGE, not the block" assertion appended straight to the page, so the two ids were equal and the claim could not fail.) Probe it: delete the line the capability rests on and watch the verify go red. If it stays green, the assertion is decoration.

The harnesses (shared, in @volter/world-core)

Vendor-agnostic; each pack feeds them its own fixtures.

  • specConformance.ts — checkSpecConformance(value, schema) → type / missing-required / enum / extra violations, plus a known-deviations allowlist and specCoverage() (implemented fields vs the spec's full set).

  • recordedDiff.ts — diffRecorded(expected, actual) → missing / extra / mismatch / type / length deviations vs a captured real response, also with a known-deviations allowlist.

  • uiConformance.ts — the UI fidelity rung (run by bun scripts/twin-conformance.ts). NOTE: this measures fidelity over what the twin models, not UI completeness — for "how much of the real product's UI exists," use the capability coverage (UI view-mode/screen capabilities). Two read-only gates:

    • Parity (checkUiCompleteness + <vendor>-ui-conformance.ts) — the mirror renders 100% of the data the twin already models (no modeled-but-unshown surface), with a test tying every rendered claim to the actual mirror bundle/state so it can't be over-claimed. This is parity between twin-data and mirror — not parity with the full product (whole real screens the twin doesn't yet model are tracked as capability todos).

    • Structural checklist (checkUiStructure + <vendor>-ui-structure.ts) — renders the mirror's components with renderToStaticMarkup and asserts the rendered DOM has the real product's structural landmarks (a column per board state, a row per item, the expected detail sections), each with an anti-vacuity teeth test.

    • Journeys (runUiJourney / browserAvailable, @volter/world-tooling's uiJourney.ts — TWIN-47/H1) — the navigability rung: seeds a throwaway twin root via the pack's real write path, boots the pack's mirror server on an ephemeral port, and drives it with REAL headless Playwright chromium using getByRole/getByText/getByLabel locators ONLY (the locators an agent transfers from the real product, never an invented className/test-id). Parity and the structural checklist are both static (renderToStaticMarkup / bundle-text greps) — neither ever boots a mirror in a browser, so a dead click handler or broken hydration passes both. Journeys close that gap: a click must actually re-render the live DOM, and selecting an item must actually flow item-specific data a list view never shows. Piloted on github (github-journey.uitest.ts). Journey and a11y rungs are named *.uitest.ts — outside the default bun test glob: they are the tests of that specific twin, run deliberately at the twin's own door or by the runner, never as freight in every ordinary suite. scripts/ui-journeys.ts is the runner twin-check.sh's [ui journeys] step calls — a REAL gate tooth when chromium is present (a broken journey turns the gate red), and a loud non-fatal advisory when the browser binary — the one non-hermetic dependency — is absent (never silently green). A negative control (uiJourney.test.ts) proves the harness has teeth: a sabotaged mirror (empty client bundle, or a button with no handler wired up) makes the journey FAIL.

      Real-app URL routing (TWIN-49/H3): all four needs-UI mirrors (github/slack/linear/jira) carry client-side routing over vendor-faithful path shapes — deep links render the linked view directly, clicks pushState (no reload), and popstate drives back/forward — and their journeys assert the pathname at each navigation step, including a dedicated deep-link

      • back/forward case per pack.

      Write journeys (TWIN-50/H4): the read journeys above only prove a mirror can be navigated; two journeys additionally prove a mirror can write through the twin's real event-sourced write path, not a UI-local mutation — slack's message composer (posting via POST /api/{method} → applySlackWrite) and linear's create-issue modal (posting a real issueCreate GraphQL mutation → executeLinearDerived). Each asserts a double: the DOM change (the new message/issue renders) AND, after the browser session tears down, the persisted twin state change — read fresh off disk, from the test process, via the same served read path the mirror itself uses. The DOM assertion alone can't catch a mirror that fakes the UI update without ever writing twin state; the persisted-state read is the anti-cheating tooth that would fail such a fake.

      The ui-scope census (each pack's census.json ui slice, checked by scripts/ui-scope.ts — TWIN-48/H2) is the committed denominator for this rung: one entry per vendor pack declaring needsUi true/false with a reason, and for the needs-UI vendors a named required-journey inventory. A pack missing from the census (or an entry for a pack that no longer exists) turns the gate RED; a required journey with no passing spec registered in JOURNEY_TEST_FILES (and no explicit todo marker) prints a loud non-fatal WARN — the visible debt line H3/H4 pick their targets from.

    Reference screenshots remain an optional, non-gating review artifact (per the "Reference screenshots" note below — never pixel-gated), added opportunistically.

Per-twin status

This table used to list only the original 5 twins (linear/slack/stripe/github/jira); the repo has since grown to 69 (this hand-maintained table currently covers 32 of them). Unlike the "Capability coverage" table below, this table is NOT generated — there is no --write-docs for the fidelity rungs yet (that generalization is still a todo, tracked as B8-follow-on). It is filled in by hand from what's actually on disk for each twin (verified per-pack: a matching harness file present, gated by a real *.test.ts that twin-check.sh runs) — treat it as best-effort and re-verify before relying on a specific cell for a specific twin.

Twin 1 derive 2 spec 3 recorded-diff 4 SDK parity 5 UI capture script
algolia — ⏳ — ✅ — —
anthropic — ✅ — — — —
aws — ✅ — ✅ ⏳ —
calcom — ✅ — ⏳ ⏳ —
clerk — ✅ — — ⏳ —
elevenlabs — ⏳ — — — —
fal — ⏳ — ✅ — —
github — ✅ ⏳ ✅ ✅ —
googlemaps — ⏳ — — — —
inngest — ⏳ — ✅ — —
jira — ✅ ⏳ ✅ ✅ —
linear ✅ (SDL) ✅ ✅ ✅ ✅ ⏳
livekit — ⏳ — — — —
mapbox — ⏳ — — — —
openai — ✅ — — — —
openrouter — ⏳ — — ⏳ —
openweather — ⏳ — — — —
pinecone — ⏳ — ✅ — —
polar — ⏳ — — — —
posthog — ✅ — ✅ ⏳ —
upstash/qstash — ⏳ — ✅ — —
replicate — ⏳ — ✅ — —
resend — ✅ — ⏳ ⏳ —
sentry — ✅ — ✅ ⏳ —
slack — ✅ ✅ ✅ ✅ ✅
stream — ✅ — ✅ — —
stripe — ✅ ⏳ ✅ ✅ —
supabase — ✅ — — ⏳ —
svix — ⏳ — ✅ — —
twilio — ⏳ — ✅ — —
vital — ⏳ — ⏳ — —
webrisk — ⏳ — — — —

✅ built + gated · ⏳ partial/not yet gated (see below) · — not applicable / not attempted. This table is FIDELITY only — a ✅ means that rung's harness is built, vendor-referenced, and gated in twin-check.sh over what the twin models. It does NOT mean the twin is complete — for completeness see the generated "Capability coverage" table below (the real "how done" number), never a percentage quoted in prose here.

What the ⏳ cells mean, per column (verified by reading each pack's harness, not just filename-matching — a lesson from this pass: a file named *-sdk.integration.test.ts does NOT always mean rung 4):

  • 2 spec ⏳: vital's harness exists (vital-conformance.ts, uses the shared checkSpecConformance) but isn't invoked by any gated test yet. elevenlabs/polar/ replicate/fal/pinecone/algolia/inngest/twilio/upstash/qstash/svix check the twin's own resource/ endpoint inventory against itself (self-referential), not against a vendor spec — even though replicate's, fal's, pinecone's, algolia's, inngest's, twilio's, upstash/qstash's, and svix's manifest/error envelopes/field shapes were themselves grounded against real fetched vendor documentation and/or a LIVE SDK trace during the build (the census.json spec slice — replicate against its first-party OpenAPI document, fal against its docs pages + the fal-js client source since fal publishes no single canonical gateway OpenAPI document, pinecone against the installed @pinecone-database/pinecone SDK's own generated TypeScript-fetch client — itself compiled from Pinecone's first-party OpenAPI document — plus targeted docs.pinecone.io reads for the handful of items the generated client doesn't settle, algolia against the installed algoliasearch SDK driven LIVE end-to-end against a throwaway local server BEFORE the handler was written, catching the batch-routed write grammar no docs skim would have, inngest against a live-fetched first-party v2 REST OpenAPI document (correcting the /v1/* build-spec guess to the real /api/v2/*) PLUS the installed inngest@4.12.0 package's own compiled source (event send, step opcodes, register target, signing algorithm), twilio against THREE live-fetched first-party OpenAPI documents (twilio_api_v2010/twilio_verify_v2/twilio_lookups_v2) PLUS the installed twilio@6.0.2 package's own compiled source (httpClient seam, signature algorithm) PLUS live-fetched public error-code reference pages — four independent sources, not just one docs skim, upstash/qstash against the ACTUALLY-INSTALLED @upstash/qstash@2.11.1 package's own compiled source — PublishToApiResponse/PublishToUrlGroupsResponse/GetLogsPayload/Log/Schedule/ UrlGroup types, PLUS a LIVE cross-SDK check of the real Receiver class against upstash/qstash's signing.ts's own JWT output before the integration test was written), svix against api.svix.com's own LIVE-fetched published OpenAPI document (id patterns/status codes/list envelope/error envelope, byte-for-byte) PLUS the installed svix@1.96.1 package's own compiled source (src/webhook.ts's Webhook class) PLUS a LIVE cross-SDK check of the real Webhook(secret).verify() accepting svix-signing.ts's PORTED (from clerk-events.ts) signature output before the integration test was written, the gated conformance harness itself is still the self-referential snapshot pattern, not a systematic per-resource spec-diff. fal's counted contract surface (checkFalConformance()'s endpointsChecked: 11) legitimately sits near this rung's conformance endpoint-count floor because fal's own real gateway surface is genuinely smaller (one queue/sync lifecycle plus webhooks/JWKS, not a broad multi-resource API) — an independent §9 skeptic confirmed all 11 counted contracts are separately implemented and separately verified, not a padded or double-counted total (TWIN-101). googlemaps/livekit/mapbox/openweather/webrisk/openrouter run ad hoc structural checks against a handful of hand-picked example requests rather than a systematic per-resource required-field spec (the pattern calcom/sentry/posthog/supabase use).
  • 4 SDK parity ⏳: calcom/resend/vital each have a file literally named *-sdk.integration.test.ts, but none of the three imports the vendor's real SDK package — they drive the twin's own server code directly, so they don't prove a real-SDK round-trip (the bar aws/github/jira/linear/posthog/replicate/sentry/slack/stream/stripe/ fal/pinecone/algolia/inngest/twilio/upstash/qstash/svix do meet, each confirmed importing the actual vendor SDK — @aws-sdk/client-s3, @octokit/rest, jira.js, @linear/sdk, posthog-node, replicate, @sentry/node, @slack/web-api, stream-chat, stripe, @fal-ai/client, @pinecone-database/pinecone, algoliasearch, inngest, twilio, @upstash/qstash, svix). upstash/qstash's real SDK exposes a genuine Client({baseUrl})/QSTASH_URL constructor override (verified live against a throwaway server before the qstash SDK integration test was written — covers publishJSON to a URL and to a urlGroup fan-out, schedules.create/get, PLUS a cross-SDK check: the real Receiver class accepts a JWT signed by upstash/qstash's signing.ts and rejects it when tampered). svix's real SDK exposes a genuine new Svix(token, {serverUrl}) constructor override (verified live against a throwaway server before svix-sdk.integration.test.ts was written — covers application/endpoint/message/eventType create + read-back through the real SDK, PLUS a cross-SDK check: the real Webhook(secret).verify() accepts a signature built by svix-signing.ts's ported buildSignedSvixDelivery and rejects it when the payload is tampered). twilio's real SDK exposes a genuine injectable httpClient constructor option (new Twilio(sid, token, {httpClient}) — verified against the installed 6.0.2 package's own BaseTwilio.js/RequestClient.js source, then live against a throwaway server, before twilio-sdk.integration.test.ts was written) — a more direct seam than fal's proxy-header trick, closer to pinecone's fetchApi/algolia's hosts transporter pattern; covers Messages create/status-poll/list, Verify start/check, Lookup fetch, and two negative paths (a real RestException round-trips the twin's error envelope). inngest's real SDK exposes a documented baseUrl/eventKey/isDev constructor override (verified live against the installed 4.12.0 package before the test was written — see inngest-sdk.integration.test.ts) — covers the Event API send path only (single/batch/idempotent); the executor/serve side (Inngest calling into a real running app) is not drivable offline by a local twin (see README). fal's real SDK has no injectable baseUrl (verified against the installed 1.10.1 package's own source) — its *-sdk.integration.test.ts instead drives the twin through the SDK's own real requestMiddleware proxy protocol (x-fal-target-url), verified working end-to-end in node against the installed package before the test was written. pinecone's real SDK exposes a more direct seam still — a documented PineconeConfiguration.fetchApi full-fetch override, shared by BOTH its control-plane and data-plane request builders — verified working end-to-end (a throwaway Node http server) before pinecone-sdk.integration.test.ts was written. algolia's real SDK exposes the most direct seam of all — a documented, TYPED hosts transporter option (no fetchApi/proxy-header trick needed) — also verified working end-to-end (a throwaway Node http server) before algolia-sdk.integration.test.ts was written; that live pass caught that saveObject/partialUpdateObject/deleteObject route through POST .../batch, not a per-object PUT/DELETE (see algolia-twin.ts's header).
  • 5 UI ⏳: a mirror UI exists (*-mirror-ui.tsx) but has no *-ui-conformance.ts / *-ui-structure.ts fidelity gate yet — only github/jira/linear/slack/stripe have both.

Rung 5 (UI) ✅ means both UI fidelity gates pass — parity (the mirror renders all data the twin models) and the structural DOM checklist. It explicitly does NOT mean the UI is complete: whole real screens/view-modes the twin doesn't model yet (e.g. Linear's list view) are tracked as capability todos, not here.

Capability coverage (generated)

The current completeness numbers, per twin — the "how done" companion to the fidelity table above.

This table is GENERATED from the capability manifests (bun scripts/twin-capabilities.ts --write-docs) — do not edit it by hand; the gate fails on drift (--check-docs). Cells are verify-proven done/total per importance tier; "out of scope" counts the explicit, reasoned carve-outs (excluded from the coverage denominator).

What the denominator measures (honest framing, TWIN-87): total is manifest-enumerated, not independently vendor-exhaustive — it counts what each pack's <vendor>-capabilities.ts declares (any status), so a vendor area that was never enumerated at all cannot appear in the count. For packs that additionally commit a top-down <VENDOR>_AREAS census of the vendor's real product areas (docs nav / OpenAPI tags) — github, openai, polar, replicate, fal, pinecone, algolia, inngest, twilio, qstash, svix, cloudflare today — a gate-wired meta-test (assertAreaCensus) further guarantees no whole area is silently missing from that denominator; other packs' totals rest on manifest authorship discipline alone until they adopt the same census.

Twin Coverage (verify-proven) Core Common Niche Out of date
webrisk 98% (65/66) 9/9 39/39 17/18 0
openrouter 97% (75/77) 13/15 43/43 19/19 0
livekit 97% (64/66) 15/15 33/34 16/17 0
googlemaps 95% (95/100) 14/14 59/62 22/24 0
moonshot 94% (73/78) 20/21 44/44 9/13 0
clerk 93% (91/98) 33/34 35/37 23/27 0
anthropic 92% (92/100) 20/21 30/30 42/49 0
posthog 91% (120/132) 21/21 63/68 36/43 0
currencyapi 90% (60/67) 10/10 50/51 0/6 0
resend 88% (68/77) 21/21 35/37 12/19 0
supabase 86% (84/98) 17/23 40/42 27/33 4
togetherai 85% (101/119) 45/46 56/60 0/13 0
calcom 85% (75/88) 18/19 34/35 23/34 0
jira 84% (120/143) 29/29 43/52 48/62 0
sentry 84% (103/122) 28/32 51/56 24/34 0
slack 82% (170/207) 28/31 68/81 74/95 0
figma 82% (143/174) 23/32 65/78 55/64 15
openai 82% (98/120) 25/27 45/49 28/44 0
oa-treasury 81% (101/124) 55/58 41/47 5/19 6
xai 77% (60/78) 21/22 30/38 9/18 0
deepinfra 77% (51/66) 31/31 17/21 3/14 0
volteridentity 75% (101/134) 40/40 44/51 17/43 0
planetscale 74% (174/236) 89/94 73/88 12/54 0
stripe 73% (155/211) 43/51 67/86 45/74 0
postmark 73% (122/167) 42/42 71/89 9/36 0
cerebras 73% (67/92) 24/24 32/46 11/22 0
ai-gateway 71% (61/86) 20/21 31/40 10/25 0
github 70% (165/236) 42/49 57/70 66/117 0
fireworks 70% (81/116) 31/33 45/60 5/23 0
stream 70% (58/83) 11/13 39/40 8/30 0
googleoauth 69% (113/164) 56/60 47/65 10/39 0
upstashvector 68% (80/117) 35/37 38/53 7/27 0
smtp 66% (105/158) 45/47 46/63 14/48 6
tavily 66% (81/123) 34/34 37/68 10/21 0
vital 66% (65/98) 14/15 25/27 26/56 0
tiktok 65% (142/218) 64/64 68/95 10/59 0
fly 65% (85/130) 48/48 32/36 5/46 0
gemini 65% (58/89) 15/18 31/34 12/37 0
mapbox 64% (82/128) 21/22 45/55 16/51 4
xidentity 63% (82/131) 45/45 28/38 9/48 3
linkedin 62% (76/122) 47/49 29/54 0/19 0
elevenlabs 62% (56/91) 26/29 30/62 — 0
mailgun 61% (123/202) 52/60 59/76 12/66 8
instagram 61% (71/116) 48/49 23/46 0/21 0
linear 60% (71/118) 14/18 30/43 27/57 3
x 60% (68/114) 47/54 21/48 0/12 8
svix 60% (45/75) 20/21 25/41 0/13 0
vercel 59% (148/250) 73/75 64/107 11/68 0
groq 59% (96/162) 22/27 58/74 16/61 6
gcs 59% (69/117) 22/22 37/49 10/46 0
openweather 59% (50/85) 14/17 30/39 6/29 0
turbopuffer 59% (48/82) 17/17 26/34 5/31 0
supermemory 59% (47/80) 16/17 26/39 5/24 0
cohere 57% (86/150) 50/62 32/57 4/31 11
perplexity 57% (79/138) 27/33 39/56 13/49 12
assemblyai 57% (49/86) 17/18 26/45 6/23 0
bluesky 56% (68/121) 48/51 18/44 2/26 0
tunnel 56% (67/119) 50/55 16/33 1/31 5
tremendous 55% (95/173) 46/49 47/86 2/38 5
upstash 55% (92/166) 47/47 28/51 17/68 0
expo 55% (47/86) 17/22 22/44 8/20 0
sendblue 55% (30/55) 16/17 14/27 0/11 5
stigg 53% (153/286) 63/69 79/158 11/59 0
langfuse 53% (118/224) 35/37 68/123 15/64 0
bitly 53% (86/163) 36/37 37/71 13/55 0
deepseek 53% (65/122) 29/36 30/47 6/39 11
pinecone 52% (38/73) 17/18 18/41 3/14 5
azure 51% (134/265) 70/84 55/110 9/71 16
inngest 51% (36/71) 21/23 15/31 0/17 5
azureformrecognizer 50% (62/123) 27/31 31/49 4/43 0
aws 49% (287/581) 117/117 135/380 35/84 0
youtube 49% (103/211) 50/55 43/68 10/88 0
tinybird 49% (95/192) 29/36 65/117 1/39 0
mistral 49% (93/188) 48/60 39/83 6/45 9
fal 48% (28/58) 10/10 18/29 0/19 3
hubspot 47% (108/228) 40/42 54/106 14/80 0
firecrawl 47% (47/99) 23/29 19/40 5/30 0
replicate 47% (31/66) 13/13 15/40 3/13 3
twelvelabs 45% (58/129) 27/33 30/62 1/34 3
discord 44% (123/278) 30/34 71/119 22/125 6
datadog 44% (98/221) 40/44 42/72 16/105 0
twilio 44% (31/70) 15/16 16/34 0/20 5
veriff 43% (72/166) 32/33 31/62 9/71 0
airtable 42% (77/183) 26/31 35/60 16/92 16
deepgram 42% (38/91) 11/13 27/45 0/33 0
polar 42% (38/91) 19/27 19/58 0/6 0
intercom 41% (84/205) 32/34 44/78 8/93 0
reddit 40% (54/136) 38/39 16/53 0/44 0
dynadot 37% (70/188) 27/35 33/74 10/79 0
segment 37% (19/52) 13/14 6/20 0/18 2
ahrefs 36% (91/252) 46/52 41/102 4/98 0
algolia 36% (26/72) 19/22 7/30 0/20 5
snowflake 35% (74/212) 41/55 25/97 8/60 16
plain 31% (149/477) 67/74 68/179 14/224 0
paypal 30% (66/220) 34/42 31/75 1/103 0
npm-registry 30% (38/128) 33/47 5/63 0/18 16
scrapecreators 28% (67/241) 29/34 26/62 12/145 0
googleads 24% (62/256) 37/50 23/116 2/90 0
axiom 18% (23/126) 11/12 12/23 0/91 0
runhuman 17% (16/92) 10/11 6/26 0/55 0
mixpanel 11% (14/123) 10/10 4/109 0/4 0
cloudflare 3% (85/3348) 14/22 65/694 6/2632 0
sendgrid 1% (3/403) 3/3 0/392 0/8 0
notion 1% (1/179) 1/45 0/78 0/56 132

Across 105 twins: 8388/18651 verify-proven done (45%) · 354 claims out of date (a pack at protocol 1: its vendor half is unproven until it moves; generated/INDEX.md names each pack's protocol).


The freshness cycle

Goal: stay close to vendor reality. The whole sweep is a re-census campaign the owner dispatches; every merge carries the fast subset. Nothing here is scheduled — this repo has no CI and nothing is cron'd (see AGENTS.md).

        ┌─────────────────────────────────────────────────────────┐
 on ask │ 1. CAPTURE   real vendor responses → recorded-diff oracle │
        │ 2. CHECK     spec + recorded-diff + SDK + UI, per vendor   │
        │ 3. REPORT    one coverage report; deviations + gaps        │
        │ 4. GATE      fail the run on regressions; gaps → worklist  │
        └─────────────────────────────────────────────────────────┘
 per-merge:  the fast subset (spec + SDK + UI completeness) for the packs a change touched
  1. Capture (refresh the oracle) — bun scripts/capture/<v>.ts. Read-only against the real vendor; rewrites the recorded-diff fixtures. A stale oracle is a false green, so a re-capture is part of any vendor re-census campaign. (Today only scripts/capture/slack.ts exists.)
  2. Check — for each vendor, run every rung it holds over a seeded or freshly-pulled world. (Conformance checks the objects present in the world, so point it at real pulled state or a seed — an empty world checks nothing.)
  3. Report — one coverage report: per-rung pass/fail, deviations, and the completeness "missing" set. This is the single artifact a human reads.
  4. Gate — a regression (new violation, dropped coverage) fails the run; new gaps become the worklist. scripts/twin-check.sh is the full sweep, run on the owner's ask; a merge clears the touched packs' own suites plus the meta-gates it affects.

Running it today

# per-vendor probe (one rung set), against a world at --root:
bun packages/twin/stripe/src/cli.ts conformance --root <world>

# the aggregate UI-rung runner (all 5 vendors, one GapReport) — wired into the gate below:
bun scripts/twin-conformance.ts --check-only

# the standing gate (all packs' conformance tests + tsc + cookbook + twin-conformance):
bash scripts/twin-check.sh

scripts/twin-conformance.ts (TWIN-99 / R-O6) is a real, wired-in gate step ("twin conformance (ui)" in scripts/twin-check.sh): it sweeps the UI-completeness + UI-structure rungs for all 5 piloted vendors into one GapReport and fails the gate on any gap NOT already recorded, with a reason, in the script's own BASELINE map — the known backlog (9 gaps today: 2 github, 1 jira, 6 slack) stays non-fatal, but a genuinely new regression turns the step red.

What is still missing (open)

  • Capture scripts for the four piloted vendors that lack one (github / stripe / jira / linear) — without them a re-census campaign has no oracle to refresh for those packs.
  • A single command that runs capture → aggregate → report for a named vendor, so a re-census campaign is one dispatch per vendor rather than a hand-assembled sequence.

Note what is deliberately NOT missing: there is no schedule to build. Freshness here is an owner-dispatched campaign, recorded as a campaign record — not a cron job nobody reads.

Reference screenshots (UI proximity)

Reference screenshots are not required for every screen. The UI-conformance run keeps a searchable reference store indexed by vendor + screen; on each run it looks one up and, when a match exists, attaches it beside the mirror's own screenshot for visual review. No match → you just get the mirror capture. References are added opportunistically (whenever a real one is captured); search is what surfaces them. Proximity is never pixel-gated — the gate is the structural checklist; the screenshots are review artifacts.

Known deviations

Every harness takes a known-deviations allowlist: a deviation we have inspected and accept (with a written reason), so the gate stays green without hiding it. A deviation must be either fixed or explicitly allowlisted with a reason — never silently ignored. The allowlist is itself reviewable evidence of what the twin does not faithfully reproduce.