# Conformance

A twin is only worth trusting if its behavior **provably tracks the real vendor**.

## Read this first — conformance answers TWO separate questions

Keep these distinct; conflating them is what produced misleading "100%" claims before.

| | **Completeness** | **Fidelity** |
|---|---|---|
| Question | *How much of the real vendor do we cover?* | *Is what we DO model faithful to the vendor?* |
| Measured by | the **capability manifest** vs the vendor's full surface | the **evidence rungs** over the modeled subset |
| Mechanism | `<vendor>-capabilities.ts` + `checkCapabilities` (a ground-truth `verify()` per capability) | spec-conformance · recorded-diff · SDK-parity · UI-structure |
| Output | honest coverage % + the ranked, categorized worklist (regression / todo by tier) | per-rung pass/fail + deviations over what's built |
| Run it | `bun scripts/twin-capabilities.ts` | `bash scripts/twin-check.sh` |

**The trap to avoid:** a twin can be *high-fidelity but low-completeness* — faithful on a
small slice. "UI rung ✅" means the mirror faithfully renders **what the twin models**; it
does **NOT** mean the UI is complete — completeness is the capability %. The honest "how
done are we" number is always the **capability coverage**, never a rung checkmark.
Faithfully here means *as the product renders it* — see [What a mirror IS](#what-a-mirror-is-read-this-before-building-one);
scope limits what the mirror shows, never how honestly it shows it.

Completeness (the headline + worklist) is described next; the fidelity rungs follow.

---

## Completeness — capability conformance (the headline metric + worklist)

See the dedicated section below ("Capability conformance"). In one line: each world's
manifest enumerates the **real vendor's full surface**; `verify()` proves each `done`;
coverage is honest and currently partial. This is the number to trust for "how complete is
the twin" — see the **generated, current** per-twin table below (B8: `bun
scripts/twin-capabilities.ts --write-docs`) rather than any percentage quoted in prose,
which is a point-in-time snapshot and goes stale the moment a manifest grows.

---

## Fidelity — the ladder of evidence rungs

These rungs prove the **quality of what the twin models** (not how much it covers —
that's completeness, above). From weakest to strongest; each is a different *kind* of
proof, not a replacement for the one below.

| # | Rung | What it proves | Oracle |
|---|------|----------------|--------|
| 1 | **Spec-derivation** | The twin's schema *cannot* drift from the vendor's, because it's generated from it | vendor SDL / OpenAPI (build-time) |
| 2 | **Spec-conformance** | Every field the twin emits matches the published spec (type / required / enum), and we know our coverage of the spec | published schema |
| 3 | **Recorded-diff** | The twin's output matches a **real captured response** byte-shape (missing / extra / mismatch / type / length) | a real vendor response on disk |
| 4 | **SDK-integration parity** | The **real vendor SDK**, unmodified, drives the twin and round-trips | the SDK itself |
| 5 | **UI conformance** | The **UI mirror** renders what the real product's UI shows (completeness) and looks structurally like it (proximity) | a declared UI surface inventory (+ optional reference screenshot) |

Rungs 1–4 cover the **API**; rung 5 is a parallel dimension covering the **UI mirror**.
A twin states which rungs it holds — see the status table below.

### What a mirror IS (read this before building one)

A mirror is **the real product's UI, running locally** — using it should feel like using the
vendor's app, minus the network. It is *not* a debug view, a data browser, or an inspector over
twin state. That is the whole reason it exists: the twin world is what the AI sees and acts on,
and the mirror is how a person **looks at that world directly** and judges whether the agent's
version of reality — and everything the agent did to it — is right. A mirror that renders a
*representation* of the data instead of the product breaks that: it looks authoritative while
showing something the vendor's UI would never show, so a reviewer signs off on a world they
never actually saw.

The concrete rule this implies: **for anything the twin models, the mirror shows the real thing,
not a stand-in.** An image renders as the image, not a labelled empty box. An avatar is the
avatar. A file thumbnail resolves. A placeholder is only acceptable where the *vendor's own UI*
shows a placeholder. "The data is present in the DOM as an attribute" is not rendering it —
`data-image-ref` on an empty div is exactly the failure this rule names, because a human
reviewing that canvas sees blank rectangles where the design's actual content lives.

This is deliberately stricter than the parity gate below, which only asks that modeled data be
*shown somewhere*. Parity catches omission; this catches **misrepresentation**, and a mirror can
pass parity while failing it.

When serving the real thing needs bytes the twin cannot synthesize (rendered images, uploaded
files, generated media), the answer is to **serve what the connector pulled** — a pulled file
carries the vendor's own assets, and the twin's job is to hand them back. Declaring the surface
a gap because the twin cannot *generate* the pixels confuses synthesis with replay: the
twin never had to invent them.

**A third, orthogonal axis: unit-test depth vs manifest-only coverage.** Completeness (above)
measures coverage of the vendor's surface; fidelity (this ladder) measures faithfulness of what's
modeled. Neither measures **how a pack proves its behavior** — a hand-authored `*.test.ts` file
vs. the capability manifest's own `verify()` round-trips (which ARE tests, and are mutation-gated
— see `scripts/mutation-test.ts` in the freshness cycle below). Raw test-file-count is a poor proxy for
this (a pack can have few `*.test.ts` files and a deep, mutation-proven manifest, or vice versa).
[./test-depth.md](./test-depth.md) is the per-pack resolution for the eight packs the
architecture review flagged as thin by that proxy — each one either gained real test depth or
carries a specific, named, written acceptance of manifest-only coverage.

## Capability conformance — the exhaustive, categorized gap check

The rung harnesses below measure a twin against its **own declarations** (a hand-kept
inventory/schema). That is **self-referential**: anything nobody declared is invisible —
which is how Linear reported "UI 100%" while missing its entire **list view mode**. The
fix is an authoritative, independent definition of *what the real vendor does*:

**The scope rule: 100% of what the vendor does.** There is no "in scope" list and no
exclusion list. Everything the real vendor does is **covered by default**, and every
capability is `done` or `todo` — nothing else. A twin is a deterministic, offline model of
the vendor's API contract; "the real thing" (live model output, real delivery, live data,
infrastructure physics, hosted pixels, real external trust) is not a capability a twin
lacks, it is the definition of a twin, so it is never listed as a gap. So every gap is a
todo — there is no "silently absent" state and no "chose not to" state. This is what stops
gaps like Linear's missing list view
from hiding — an un-enumerated, not-carved-out capability is, by definition, an open todo.

- Each world declares a **capability manifest** (`<vendor>-capabilities.ts`) — the
  expected real-product surface: every core capability across **API + UI view-modes +
  connector**. This is "what 100% means," reviewed and separate from the implementation.
- Each capability carries a ground-truth **`verify()`** predicate (UI: the built mirror
  renders the screen; API: the operation runs), so **`done` must be proven, not asserted.**
- Each capability also carries an importance **`tier`** — `core` (the daily-use backbone),
  `common` (frequently used), or `niche` (long-tail / admin / power-user). The tier does
  **not** change scope (everything is built or a todo); it only **ranks the worklist**
  so we close the highest-value gaps first.
- `checkCapabilities` (in `@volter/world-tooling` — `packages/world-tooling/src/capabilities.ts`,
  dev-only, NOT the runtime kernel `@volter/world-core`) reconciles claim-vs-truth into the exhaustive,
  categorized, **ranked** worklist (`bun scripts/twin-capabilities.ts`):
  - **regression** — declared done but `verify()` now fails (a false-green / breakage). Always first.
  - **todo** — covered by default, not built yet — the worklist, **ordered core → common → niche**.
  - *(done — expected and `verify()` passes.)*
- A world with **no manifest yet is itself the biggest todo** (declare it). Each cycle
  grows every manifest toward the real product's full surface, then closes regression +
  todo (**core todos first**). The check **produces the ranked list**; the closer (subagent)
  **closes it from the top**; re-running proves closure. All 69 twins now carry a manifest
  against their real-vendor surface (honest, verify-proven coverage) — for the **current**
  per-twin numbers and total todos, see the generated table below
  ("Capability coverage") or run `bun scripts/twin-capabilities.ts`. *(Historical snapshot,
  now stale: at the original 5-world baseline — before this repo grew to 26 twins — coverage
  read slack 50% (75/150), stripe 46% (58/126), linear 44% (26/59), jira 42% (49/118), github
  36% (38/107), 314 todos. Numbers drift every cycle; never treat a prose
  percentage as current — the generated table is the only one the drift gate keeps honest.)*

### Adding a twin — conformance checklist (and the anti-patterns that bit us)

A new twin is not done until it satisfies these. The meta-test
(`scripts/capability-manifests.test.ts`, run by `twin-check`) **enforces** items 1, 4, 5, 6
mechanically; items 2–3 are judgment and are guarded by adversarial review each cycle.

1. **Declare `<vendor>-capabilities.ts`** — a `CapabilitySpec[]` + `<vendor>Capabilities()`
   calling `checkCapabilities`. Wire it into `scripts/capability-manifests.ts`.
2. **The manifest is the REAL vendor's full surface (the target), authored top-down —
   NOT a list of what you built.** A manifest that mirrors your implementation makes the
   metric meaningless (a thin list inflates coverage). Most entries start `todo`; coverage
   should read honestly LOW. *(Anti-pattern that bit us: the first manifests were
   self-portraits → "75%/100%"; rewritten to the real surface → an honest ~30–40%.)*
3. **`done` requires a ground-truth `verify()` that is a REAL round-trip** — create/mutate
   via the twin → read it back → assert the returned **values**; negative cases assert the
   vendor-shaped error. It MUST be *failable* (would fail if the capability broke). **Never**
   verify with "no errors on an empty workspace" / "an empty list parses" / a bundle string
   present for unrelated reasons. *(Anti-pattern that bit us: Linear's `done`s passed against
   an empty workspace and proved nothing; Stripe's `webhooks.emit` only checked a create
   succeeded. Both were false-greens caught only by adversarial review.)* Reference the
   strong patterns in `stripe-capabilities.ts` and `linear-capabilities.ts` (`withRoot` +
   create-read-assert).
4. **Enumerate UI VIEW-MODES and screens, not just data fields** — list/board/timeline/
   calendar, detail panes, inbox, settings, etc. A feature is `done` only when it's in the
   API **and** rendered in the UI mirror (parity). *(Anti-pattern: Linear was "UI 100%"
   while missing its entire list view — nobody had enumerated it.)*
5. **Everything the vendor does, ranked by tier.** Categories are `regression` / `todo` /
   `done`. There is no "in scope" list and no exclusion list — scope is **100% of what the
   vendor does** — every gap is a todo, never silently absent and never "chosen away".
   Tag every capability with a **`tier`**
   (`core`/`common`/`niche`) so the worklist ranks — build core gaps first. Don't report
   "deep ✅ / 100%" — report the honest measured % against the real surface.
6. **Add `<vendor>-capabilities.test.ts`** asserting `assertManifestBaseline(report)`
   (0 regressions, done>0, total>=50).
7. **Assert what DISTINGUISHES the behaviour, not merely that it failed.** A `verify()` that
   checks only `400 validation_error` proves almost nothing on an endpoint with more than one
   way to reject: delete the guard under test and a different check further down returns the
   same status, so the capability stays green while the behaviour is gone. Pin the vendor's
   **message** (`isErrSaying`) whenever several rejections share a code. The same trap has a
   positive form — a fixture where two values coincide. *(Anti-patterns, all caught by
   sabotage-probing rather than by review: Notion's read-only-property, title-removal and
   version-mismatch guards could each be deleted with every test still passing, because a
   second path produced the same 400; and its "a content webhook names the PAGE, not the
   block" assertion appended straight to the page, so the two ids were equal and the claim
   could not fail.)* **Probe it: delete the line the capability rests on and watch the verify
   go red.** If it stays green, the assertion is decoration.

## The harnesses (shared, in `@volter/world-core`)

Vendor-agnostic; each pack feeds them its own fixtures.

- **`specConformance.ts`** — `checkSpecConformance(value, schema)` → `type` /
  `missing-required` / `enum` / `extra` violations, plus a **known-deviations**
  allowlist and `specCoverage()` (implemented fields vs the spec's full set).
- **`recordedDiff.ts`** — `diffRecorded(expected, actual)` → `missing` / `extra` /
  `mismatch` / `type` / `length` deviations vs a captured real response, also with
  a known-deviations allowlist.
- **`uiConformance.ts`** — the UI **fidelity** rung (run by `bun scripts/twin-conformance.ts`).
  NOTE: this measures fidelity over *what the twin models*, **not** UI completeness — for
  "how much of the real product's UI exists," use the capability coverage (UI view-mode/screen
  capabilities). Two read-only gates:
  - **Parity** (`checkUiCompleteness` + `<vendor>-ui-conformance.ts`) — the mirror renders
    100% of the data the twin **already models** (no `modeled-but-unshown` surface), with a
    test tying every `rendered` claim to the actual mirror bundle/state so it can't be
    over-claimed. This is parity between twin-data and mirror — *not* parity with the full
    product (whole real screens the twin doesn't yet model are tracked as capability todos).
  - **Structural checklist** (`checkUiStructure` + `<vendor>-ui-structure.ts`) — renders the
    mirror's components with `renderToStaticMarkup` and asserts the rendered DOM has the
    real product's structural landmarks (a column per board state, a row per item, the
    expected detail sections), each with an anti-vacuity teeth test.
  - **Journeys** (`runUiJourney` / `browserAvailable`, `@volter/world-tooling`'s `uiJourney.ts`
    — TWIN-47/H1) — the navigability rung: seeds a throwaway twin root via the pack's real
    write path, boots the pack's mirror server on an ephemeral port, and drives it with REAL
    headless Playwright chromium using `getByRole`/`getByText`/`getByLabel` locators ONLY (the
    locators an agent transfers from the real product, never an invented className/test-id).
    Parity and the structural checklist are both static (`renderToStaticMarkup` / bundle-text
    greps) — neither ever boots a mirror in a browser, so a dead click handler or broken
    hydration passes both. Journeys close that gap: a click must actually re-render the live
    DOM, and selecting an item must actually flow item-specific data a list view never shows.
    Piloted on **github** (`github-journey.uitest.ts`). Journey and a11y rungs are named
    `*.uitest.ts` — outside the default `bun test` glob: they are the tests of that specific
    twin, run deliberately at the twin's own door or by the runner, never as freight in every
    ordinary suite. `scripts/ui-journeys.ts` is the runner
    `twin-check.sh`'s `[ui journeys]` step calls — a REAL gate tooth when chromium is present
    (a broken journey turns the gate red), and a loud non-fatal advisory when the browser
    binary — the one non-hermetic dependency — is absent (never silently green). A negative
    control (`uiJourney.test.ts`) proves the harness has teeth: a sabotaged mirror (empty
    client bundle, or a button with no handler wired up) makes the journey FAIL.

    **Real-app URL routing (TWIN-49/H3):** all four needs-UI mirrors (github/slack/linear/jira)
    carry client-side routing over vendor-faithful path shapes — deep links render the linked
    view directly, clicks `pushState` (no reload), and `popstate` drives back/forward — and
    their journeys assert the pathname at each navigation step, including a dedicated deep-link
    + back/forward case per pack.

    **Write journeys (TWIN-50/H4):** the read journeys above only prove a mirror can be
    navigated; two journeys additionally prove a mirror can **write** through the twin's real
    event-sourced write path, not a UI-local mutation — slack's message composer (posting via
    `POST /api/{method}` → `applySlackWrite`) and linear's create-issue modal (posting a real
    `issueCreate` GraphQL mutation → `executeLinearDerived`). Each asserts a **double**: the DOM
    change (the new message/issue renders) AND, after the browser session tears down, the
    persisted twin state change — read fresh off disk, from the test process, via the same
    served read path the mirror itself uses. The DOM assertion alone can't catch a mirror that
    fakes the UI update without ever writing twin state; the persisted-state read is the
    anti-cheating tooth that would fail such a fake.

    The **ui-scope census** (each pack's census.json `ui` slice, checked by `scripts/ui-scope.ts` — TWIN-48/H2) is the committed
    denominator for this rung: one entry per vendor pack declaring `needsUi` true/false with a
    reason, and for the needs-UI vendors a named required-journey inventory. A pack missing from
    the census (or an entry for a pack that no longer exists) turns the gate RED; a required
    journey with no passing spec registered in `JOURNEY_TEST_FILES` (and no explicit `todo`
    marker) prints a loud non-fatal WARN — the visible debt line H3/H4 pick their targets from.

  Reference **screenshots** remain an optional, non-gating review artifact (per the
  "Reference screenshots" note below — never pixel-gated), added opportunistically.

## Per-twin status

This table used to list only the original 5 twins (linear/slack/stripe/github/jira); the repo
has since grown to 69 (this hand-maintained table currently covers 32 of them). Unlike
the "Capability coverage" table below, **this table is NOT generated** — there is no
`--write-docs` for the fidelity rungs yet (that generalization is still a todo, tracked
as B8-follow-on). It is filled in by hand from what's actually on disk
for each twin (verified per-pack: a matching harness file present, gated by a real
`*.test.ts` that `twin-check.sh` runs) — treat it as best-effort and re-verify before relying
on a specific cell for a specific twin.

| Twin | 1 derive | 2 spec | 3 recorded-diff | 4 SDK parity | 5 UI | capture script |
|------|:--:|:--:|:--:|:--:|:--:|:--:|
| algolia | — | ⏳ | — | ✅ | — | — |
| anthropic | — | ✅ | — | — | — | — |
| aws | — | ✅ | — | ✅ | ⏳ | — |
| calcom | — | ✅ | — | ⏳ | ⏳ | — |
| clerk | — | ✅ | — | — | ⏳ | — |
| elevenlabs | — | ⏳ | — | — | — | — |
| fal | — | ⏳ | — | ✅ | — | — |
| github | — | ✅ | ⏳ | ✅ | ✅ | — |
| googlemaps | — | ⏳ | — | — | — | — |
| inngest | — | ⏳ | — | ✅ | — | — |
| jira | — | ✅ | ⏳ | ✅ | ✅ | — |
| linear | ✅ (SDL) | ✅ | ✅ | ✅ | ✅ | ⏳ |
| livekit | — | ⏳ | — | — | — | — |
| mapbox | — | ⏳ | — | — | — | — |
| openai | — | ✅ | — | — | — | — |
| openrouter | — | ⏳ | — | — | ⏳ | — |
| openweather | — | ⏳ | — | — | — | — |
| pinecone | — | ⏳ | — | ✅ | — | — |
| polar | — | ⏳ | — | — | — | — |
| posthog | — | ✅ | — | ✅ | ⏳ | — |
| upstash/qstash | — | ⏳ | — | ✅ | — | — |
| replicate | — | ⏳ | — | ✅ | — | — |
| resend | — | ✅ | — | ⏳ | ⏳ | — |
| sentry | — | ✅ | — | ✅ | ⏳ | — |
| slack | — | ✅ | ✅ | ✅ | ✅ | ✅ |
| stream | — | ✅ | — | ✅ | — | — |
| stripe | — | ✅ | ⏳ | ✅ | ✅ | — |
| supabase | — | ✅ | — | — | ⏳ | — |
| svix | — | ⏳ | — | ✅ | — | — |
| twilio | — | ⏳ | — | ✅ | — | — |
| vital | — | ⏳ | — | ⏳ | — | — |
| webrisk | — | ⏳ | — | — | — | — |

✅ built + gated · ⏳ partial/not yet gated (see below) · — not applicable / not attempted.
**This table is FIDELITY only** — a ✅ means that rung's harness is built, vendor-referenced,
and gated in `twin-check.sh` over what the twin models. It does **NOT** mean the twin is
complete — for completeness see the **generated** "Capability coverage" table below (the real
"how done" number), never a percentage quoted in prose here.

What the `⏳` cells mean, per column (verified by reading each pack's harness, not just
filename-matching — a lesson from this pass: a file named `*-sdk.integration.test.ts` does
NOT always mean rung 4):
- **2 spec** `⏳`: `vital`'s harness exists (`vital-conformance.ts`, uses the shared
  `checkSpecConformance`) but isn't invoked by any gated test yet. `elevenlabs`/`polar`/
  `replicate`/`fal`/`pinecone`/`algolia`/`inngest`/`twilio`/`upstash/qstash`/`svix` check the twin's own
  resource/ endpoint inventory against itself (self-referential), not against a vendor spec — even
  though `replicate`'s, `fal`'s, `pinecone`'s, `algolia`'s, `inngest`'s, `twilio`'s, `upstash/qstash`'s, and
  `svix`'s manifest/error envelopes/field shapes were themselves grounded against real fetched
  vendor documentation and/or
  a LIVE SDK trace during the build (the census.json `spec` slice — `replicate` against its first-party
  OpenAPI document, `fal` against its docs pages + the fal-js client source since fal publishes no
  single canonical gateway OpenAPI document, `pinecone` against the installed
  `@pinecone-database/pinecone` SDK's own generated TypeScript-fetch client — itself compiled from
  Pinecone's first-party OpenAPI document — plus targeted docs.pinecone.io reads for the handful
  of items the generated client doesn't settle, `algolia` against the installed `algoliasearch`
  SDK driven LIVE end-to-end against a throwaway local server BEFORE the handler was written,
  catching the batch-routed write grammar no docs skim would have, `inngest` against a
  live-fetched first-party v2 REST OpenAPI document (correcting the `/v1/*` build-spec guess to
  the real `/api/v2/*`) PLUS the installed `inngest@4.12.0` package's own compiled source (event
  send, step opcodes, register target, signing algorithm), `twilio` against THREE live-fetched
  first-party OpenAPI documents (`twilio_api_v2010`/`twilio_verify_v2`/`twilio_lookups_v2`) PLUS
  the installed `twilio@6.0.2` package's own compiled source (httpClient seam, signature algorithm)
  PLUS live-fetched public error-code reference pages — four independent sources, not just one docs
  skim, `upstash/qstash` against the ACTUALLY-INSTALLED `@upstash/qstash@2.11.1` package's own compiled
  source — `PublishToApiResponse`/`PublishToUrlGroupsResponse`/`GetLogsPayload`/`Log`/`Schedule`/
  `UrlGroup` types, PLUS a LIVE cross-SDK check of the real `Receiver` class against
  `upstash/qstash`'s `signing.ts`'s own JWT output before the integration test was written), `svix` against
  `api.svix.com`'s own LIVE-fetched published OpenAPI document (id patterns/status codes/list
  envelope/error envelope, byte-for-byte) PLUS the installed `svix@1.96.1` package's own compiled
  source (`src/webhook.ts`'s `Webhook` class) PLUS a LIVE cross-SDK check of the real
  `Webhook(secret).verify()` accepting `svix-signing.ts`'s PORTED (from `clerk-events.ts`) signature
  output before the integration test was written, the gated
  conformance harness itself is still the self-referential snapshot
  pattern, not a systematic per-resource spec-diff. `fal`'s counted
  contract surface (`checkFalConformance()`'s `endpointsChecked: 11`) legitimately sits near this
  rung's conformance endpoint-count floor because fal's own real gateway surface is genuinely
  smaller (one queue/sync lifecycle plus webhooks/JWKS, not a broad multi-resource API) — an
  independent §9 skeptic confirmed all 11 counted contracts are separately implemented and
  separately verified, not a padded or double-counted total (TWIN-101).
  `googlemaps`/`livekit`/`mapbox`/`openweather`/`webrisk`/`openrouter` run ad hoc
  structural checks against a handful of hand-picked example requests rather than a systematic
  per-resource required-field spec (the pattern `calcom`/`sentry`/`posthog`/`supabase` use).
- **4 SDK parity** `⏳`: `calcom`/`resend`/`vital` each have a file literally named
  `*-sdk.integration.test.ts`, but none of the three imports the vendor's real SDK package —
  they drive the twin's own server code directly, so they don't prove a real-SDK round-trip
  (the bar `aws`/`github`/`jira`/`linear`/`posthog`/`replicate`/`sentry`/`slack`/`stream`/`stripe`/
  `fal`/`pinecone`/`algolia`/`inngest`/`twilio`/`upstash/qstash`/`svix` do meet, each confirmed importing the actual
  vendor SDK — `@aws-sdk/client-s3`, `@octokit/rest`, `jira.js`, `@linear/sdk`, `posthog-node`, `replicate`,
  `@sentry/node`, `@slack/web-api`, `stream-chat`, `stripe`, `@fal-ai/client`,
  `@pinecone-database/pinecone`, `algoliasearch`, `inngest`, `twilio`, `@upstash/qstash`, `svix`). `upstash/qstash`'s
  real SDK exposes a genuine `Client({baseUrl})`/`QSTASH_URL` constructor override (verified live
  against a throwaway server before the qstash SDK integration test was written — covers
  `publishJSON` to a URL and to a urlGroup fan-out, `schedules.create`/`get`, PLUS a cross-SDK
  check: the real `Receiver` class accepts a JWT signed by `upstash/qstash`'s `signing.ts` and rejects it when
  tampered). `svix`'s real SDK exposes a genuine `new Svix(token, {serverUrl})` constructor
  override (verified live against a throwaway server before `svix-sdk.integration.test.ts` was
  written — covers application/endpoint/message/eventType create + read-back through the real
  SDK, PLUS a cross-SDK check: the real `Webhook(secret).verify()` accepts a signature built by
  `svix-signing.ts`'s ported `buildSignedSvixDelivery` and rejects it when the payload is
  tampered). `twilio`'s real SDK exposes
  a genuine injectable `httpClient` constructor option (`new Twilio(sid, token, {httpClient})` —
  verified against the installed 6.0.2 package's own `BaseTwilio.js`/`RequestClient.js` source,
  then live against a throwaway server, before `twilio-sdk.integration.test.ts` was written) — a
  more direct seam than `fal`'s proxy-header trick, closer to `pinecone`'s `fetchApi`/`algolia`'s
  `hosts` transporter pattern; covers Messages create/status-poll/list, Verify start/check, Lookup
  fetch, and two negative paths (a real `RestException` round-trips the twin's error envelope).
  `inngest`'s real SDK exposes a
  documented `baseUrl`/`eventKey`/`isDev` constructor override (verified live against the
  installed 4.12.0 package before the test was written — see inngest-sdk.integration.test.ts) —
  covers the Event API send path only (single/batch/idempotent); the executor/serve side (Inngest
  calling into a real running app) is not drivable offline by a local twin (see README).
  `fal`'s real SDK has no injectable `baseUrl` (verified against the installed 1.10.1 package's
  own source) — its `*-sdk.integration.test.ts` instead drives the twin through the SDK's own real
  `requestMiddleware` proxy protocol (`x-fal-target-url`), verified working end-to-end in node
  against the installed package before the test was written. `pinecone`'s real SDK exposes a more
  direct seam still — a documented `PineconeConfiguration.fetchApi` full-fetch override, shared by
  BOTH its control-plane and data-plane request builders — verified working end-to-end (a
  throwaway Node http server) before `pinecone-sdk.integration.test.ts` was written. `algolia`'s
  real SDK exposes the most direct seam of all — a documented, TYPED `hosts` transporter option
  (no fetchApi/proxy-header trick needed) — also verified working end-to-end (a throwaway Node
  http server) before `algolia-sdk.integration.test.ts` was written; that live pass caught that
  `saveObject`/`partialUpdateObject`/`deleteObject` route through `POST .../batch`, not a
  per-object PUT/DELETE (see algolia-twin.ts's header).
- **5 UI** `⏳`: a mirror UI exists (`*-mirror-ui.tsx`) but has no `*-ui-conformance.ts` /
  `*-ui-structure.ts` fidelity gate yet — only github/jira/linear/slack/stripe have both.

**Rung 5 (UI) ✅** means both UI fidelity gates pass — parity (the mirror renders all data
the twin models) **and** the structural DOM checklist. It explicitly does NOT mean the UI is
complete: whole real screens/view-modes the twin doesn't model yet (e.g. Linear's list view)
are tracked as **capability todos**, not here.

## Capability coverage (generated)

The current completeness numbers, per twin — the "how done" companion to the fidelity table above.

<!-- capability-coverage:BEGIN — GENERATED by `bun scripts/twin-capabilities.ts --write-docs`; do not edit between markers (`--check-docs` gates drift) -->

*This table is **GENERATED** from the capability manifests (`bun scripts/twin-capabilities.ts --write-docs`) — do not edit it by hand; the gate fails on drift (`--check-docs`). Cells are verify-proven **done/total** per importance tier; "out of scope" counts the explicit, reasoned carve-outs (excluded from the coverage denominator).*

*What the denominator measures (honest framing, TWIN-87): `total` is **manifest-enumerated**, not independently vendor-exhaustive — it counts what each pack's `<vendor>-capabilities.ts` declares (any status), so a vendor area that was never enumerated at all cannot appear in the count. For packs that additionally commit a top-down `<VENDOR>_AREAS` census of the vendor's real product areas (docs nav / OpenAPI tags) — github, openai, polar, replicate, fal, pinecone, algolia, inngest, twilio, qstash, svix, cloudflare today — a gate-wired meta-test (`assertAreaCensus`) further guarantees no whole area is silently missing from that denominator; other packs' totals rest on manifest authorship discipline alone until they adopt the same census.*

| Twin | Coverage (verify-proven) | Core | Common | Niche | Out of date |
|------|:--:|:--:|:--:|:--:|:--:|
| [webrisk](../../packages/twin/webrisk) | **98%** (65/66) | 9/9 | 39/39 | 17/18 | 0 |
| [openrouter](../../packages/twin/openrouter) | **97%** (75/77) | 13/15 | 43/43 | 19/19 | 0 |
| [livekit](../../packages/twin/livekit) | **97%** (64/66) | 15/15 | 33/34 | 16/17 | 0 |
| [googlemaps](../../packages/twin/googlemaps) | **95%** (95/100) | 14/14 | 59/62 | 22/24 | 0 |
| [moonshot](../../packages/twin/moonshot) | **94%** (73/78) | 20/21 | 44/44 | 9/13 | 0 |
| [clerk](../../packages/twin/clerk) | **93%** (91/98) | 33/34 | 35/37 | 23/27 | 0 |
| [anthropic](../../packages/twin/anthropic) | **92%** (92/100) | 20/21 | 30/30 | 42/49 | 0 |
| [posthog](../../packages/twin/posthog) | **91%** (120/132) | 21/21 | 63/68 | 36/43 | 0 |
| [currencyapi](../../packages/twin/currencyapi) | **90%** (60/67) | 10/10 | 50/51 | 0/6 | 0 |
| [resend](../../packages/twin/resend) | **88%** (68/77) | 21/21 | 35/37 | 12/19 | 0 |
| [supabase](../../packages/twin/supabase) | **86%** (84/98) | 17/23 | 40/42 | 27/33 | 4 |
| [togetherai](../../packages/twin/togetherai) | **85%** (101/119) | 45/46 | 56/60 | 0/13 | 0 |
| [calcom](../../packages/twin/calcom) | **85%** (75/88) | 18/19 | 34/35 | 23/34 | 0 |
| [jira](../../packages/twin/jira) | **84%** (120/143) | 29/29 | 43/52 | 48/62 | 0 |
| [sentry](../../packages/twin/sentry) | **84%** (103/122) | 28/32 | 51/56 | 24/34 | 0 |
| [slack](../../packages/twin/slack) | **82%** (170/207) | 28/31 | 68/81 | 74/95 | 0 |
| [figma](../../packages/twin/figma) | **82%** (143/174) | 23/32 | 65/78 | 55/64 | 15 |
| [openai](../../packages/twin/openai) | **82%** (98/120) | 25/27 | 45/49 | 28/44 | 0 |
| [oa-treasury](../../packages/twin/oa-treasury) | **81%** (101/124) | 55/58 | 41/47 | 5/19 | 6 |
| [xai](../../packages/twin/xai) | **77%** (60/78) | 21/22 | 30/38 | 9/18 | 0 |
| [deepinfra](../../packages/twin/deepinfra) | **77%** (51/66) | 31/31 | 17/21 | 3/14 | 0 |
| [volteridentity](../../packages/twin/volteridentity) | **75%** (101/134) | 40/40 | 44/51 | 17/43 | 0 |
| [planetscale](../../packages/twin/planetscale) | **74%** (174/236) | 89/94 | 73/88 | 12/54 | 0 |
| [stripe](../../packages/twin/stripe) | **73%** (155/211) | 43/51 | 67/86 | 45/74 | 0 |
| [postmark](../../packages/twin/postmark) | **73%** (122/167) | 42/42 | 71/89 | 9/36 | 0 |
| [cerebras](../../packages/twin/cerebras) | **73%** (67/92) | 24/24 | 32/46 | 11/22 | 0 |
| [ai-gateway](../../packages/twin/ai-gateway) | **71%** (61/86) | 20/21 | 31/40 | 10/25 | 0 |
| [github](../../packages/twin/github) | **70%** (165/236) | 42/49 | 57/70 | 66/117 | 0 |
| [fireworks](../../packages/twin/fireworks) | **70%** (81/116) | 31/33 | 45/60 | 5/23 | 0 |
| [stream](../../packages/twin/stream) | **70%** (58/83) | 11/13 | 39/40 | 8/30 | 0 |
| [googleoauth](../../packages/twin/googleoauth) | **69%** (113/164) | 56/60 | 47/65 | 10/39 | 0 |
| [upstashvector](../../packages/twin/upstashvector) | **68%** (80/117) | 35/37 | 38/53 | 7/27 | 0 |
| [smtp](../../packages/twin/smtp) | **66%** (105/158) | 45/47 | 46/63 | 14/48 | 6 |
| [tavily](../../packages/twin/tavily) | **66%** (81/123) | 34/34 | 37/68 | 10/21 | 0 |
| [vital](../../packages/twin/vital) | **66%** (65/98) | 14/15 | 25/27 | 26/56 | 0 |
| [tiktok](../../packages/twin/tiktok) | **65%** (142/218) | 64/64 | 68/95 | 10/59 | 0 |
| [fly](../../packages/twin/fly) | **65%** (85/130) | 48/48 | 32/36 | 5/46 | 0 |
| [gemini](../../packages/twin/gemini) | **65%** (58/89) | 15/18 | 31/34 | 12/37 | 0 |
| [mapbox](../../packages/twin/mapbox) | **64%** (82/128) | 21/22 | 45/55 | 16/51 | 4 |
| [xidentity](../../packages/twin/xidentity) | **63%** (82/131) | 45/45 | 28/38 | 9/48 | 3 |
| [linkedin](../../packages/twin/linkedin) | **62%** (76/122) | 47/49 | 29/54 | 0/19 | 0 |
| [elevenlabs](../../packages/twin/elevenlabs) | **62%** (56/91) | 26/29 | 30/62 | — | 0 |
| [mailgun](../../packages/twin/mailgun) | **61%** (123/202) | 52/60 | 59/76 | 12/66 | 8 |
| [instagram](../../packages/twin/instagram) | **61%** (71/116) | 48/49 | 23/46 | 0/21 | 0 |
| [linear](../../packages/twin/linear) | **60%** (71/118) | 14/18 | 30/43 | 27/57 | 3 |
| [x](../../packages/twin/x) | **60%** (68/114) | 47/54 | 21/48 | 0/12 | 8 |
| [svix](../../packages/twin/svix) | **60%** (45/75) | 20/21 | 25/41 | 0/13 | 0 |
| [vercel](../../packages/twin/vercel) | **59%** (148/250) | 73/75 | 64/107 | 11/68 | 0 |
| [groq](../../packages/twin/groq) | **59%** (96/162) | 22/27 | 58/74 | 16/61 | 6 |
| [gcs](../../packages/twin/gcs) | **59%** (69/117) | 22/22 | 37/49 | 10/46 | 0 |
| [openweather](../../packages/twin/openweather) | **59%** (50/85) | 14/17 | 30/39 | 6/29 | 0 |
| [turbopuffer](../../packages/twin/turbopuffer) | **59%** (48/82) | 17/17 | 26/34 | 5/31 | 0 |
| [supermemory](../../packages/twin/supermemory) | **59%** (47/80) | 16/17 | 26/39 | 5/24 | 0 |
| [cohere](../../packages/twin/cohere) | **57%** (86/150) | 50/62 | 32/57 | 4/31 | 11 |
| [perplexity](../../packages/twin/perplexity) | **57%** (79/138) | 27/33 | 39/56 | 13/49 | 12 |
| [assemblyai](../../packages/twin/assemblyai) | **57%** (49/86) | 17/18 | 26/45 | 6/23 | 0 |
| [bluesky](../../packages/twin/bluesky) | **56%** (68/121) | 48/51 | 18/44 | 2/26 | 0 |
| [tunnel](../../packages/twin/tunnel) | **56%** (67/119) | 50/55 | 16/33 | 1/31 | 5 |
| [tremendous](../../packages/twin/tremendous) | **55%** (95/173) | 46/49 | 47/86 | 2/38 | 5 |
| [upstash](../../packages/twin/upstash) | **55%** (92/166) | 47/47 | 28/51 | 17/68 | 0 |
| [expo](../../packages/twin/expo) | **55%** (47/86) | 17/22 | 22/44 | 8/20 | 0 |
| [sendblue](../../packages/twin/sendblue) | **55%** (30/55) | 16/17 | 14/27 | 0/11 | 5 |
| [stigg](../../packages/twin/stigg) | **53%** (153/286) | 63/69 | 79/158 | 11/59 | 0 |
| [langfuse](../../packages/twin/langfuse) | **53%** (118/224) | 35/37 | 68/123 | 15/64 | 0 |
| [bitly](../../packages/twin/bitly) | **53%** (86/163) | 36/37 | 37/71 | 13/55 | 0 |
| [deepseek](../../packages/twin/deepseek) | **53%** (65/122) | 29/36 | 30/47 | 6/39 | 11 |
| [pinecone](../../packages/twin/pinecone) | **52%** (38/73) | 17/18 | 18/41 | 3/14 | 5 |
| [azure](../../packages/twin/azure) | **51%** (134/265) | 70/84 | 55/110 | 9/71 | 16 |
| [inngest](../../packages/twin/inngest) | **51%** (36/71) | 21/23 | 15/31 | 0/17 | 5 |
| [azureformrecognizer](../../packages/twin/azureformrecognizer) | **50%** (62/123) | 27/31 | 31/49 | 4/43 | 0 |
| [aws](../../packages/twin/aws) | **49%** (287/581) | 117/117 | 135/380 | 35/84 | 0 |
| [youtube](../../packages/twin/youtube) | **49%** (103/211) | 50/55 | 43/68 | 10/88 | 0 |
| [tinybird](../../packages/twin/tinybird) | **49%** (95/192) | 29/36 | 65/117 | 1/39 | 0 |
| [mistral](../../packages/twin/mistral) | **49%** (93/188) | 48/60 | 39/83 | 6/45 | 9 |
| [fal](../../packages/twin/fal) | **48%** (28/58) | 10/10 | 18/29 | 0/19 | 3 |
| [hubspot](../../packages/twin/hubspot) | **47%** (108/228) | 40/42 | 54/106 | 14/80 | 0 |
| [firecrawl](../../packages/twin/firecrawl) | **47%** (47/99) | 23/29 | 19/40 | 5/30 | 0 |
| [replicate](../../packages/twin/replicate) | **47%** (31/66) | 13/13 | 15/40 | 3/13 | 3 |
| [twelvelabs](../../packages/twin/twelvelabs) | **45%** (58/129) | 27/33 | 30/62 | 1/34 | 3 |
| [discord](../../packages/twin/discord) | **44%** (123/278) | 30/34 | 71/119 | 22/125 | 6 |
| [datadog](../../packages/twin/datadog) | **44%** (98/221) | 40/44 | 42/72 | 16/105 | 0 |
| [twilio](../../packages/twin/twilio) | **44%** (31/70) | 15/16 | 16/34 | 0/20 | 5 |
| [veriff](../../packages/twin/veriff) | **43%** (72/166) | 32/33 | 31/62 | 9/71 | 0 |
| [airtable](../../packages/twin/airtable) | **42%** (77/183) | 26/31 | 35/60 | 16/92 | 16 |
| [deepgram](../../packages/twin/deepgram) | **42%** (38/91) | 11/13 | 27/45 | 0/33 | 0 |
| [polar](../../packages/twin/polar) | **42%** (38/91) | 19/27 | 19/58 | 0/6 | 0 |
| [intercom](../../packages/twin/intercom) | **41%** (84/205) | 32/34 | 44/78 | 8/93 | 0 |
| [reddit](../../packages/twin/reddit) | **40%** (54/136) | 38/39 | 16/53 | 0/44 | 0 |
| [dynadot](../../packages/twin/dynadot) | **37%** (70/188) | 27/35 | 33/74 | 10/79 | 0 |
| [segment](../../packages/twin/segment) | **37%** (19/52) | 13/14 | 6/20 | 0/18 | 2 |
| [ahrefs](../../packages/twin/ahrefs) | **36%** (91/252) | 46/52 | 41/102 | 4/98 | 0 |
| [algolia](../../packages/twin/algolia) | **36%** (26/72) | 19/22 | 7/30 | 0/20 | 5 |
| [snowflake](../../packages/twin/snowflake) | **35%** (74/212) | 41/55 | 25/97 | 8/60 | 16 |
| [plain](../../packages/twin/plain) | **31%** (149/477) | 67/74 | 68/179 | 14/224 | 0 |
| [paypal](../../packages/twin/paypal) | **30%** (66/220) | 34/42 | 31/75 | 1/103 | 0 |
| [npm-registry](../../packages/twin/npm-registry) | **30%** (38/128) | 33/47 | 5/63 | 0/18 | 16 |
| [scrapecreators](../../packages/twin/scrapecreators) | **28%** (67/241) | 29/34 | 26/62 | 12/145 | 0 |
| [googleads](../../packages/twin/googleads) | **24%** (62/256) | 37/50 | 23/116 | 2/90 | 0 |
| [axiom](../../packages/twin/axiom) | **18%** (23/126) | 11/12 | 12/23 | 0/91 | 0 |
| [runhuman](../../packages/twin/runhuman) | **17%** (16/92) | 10/11 | 6/26 | 0/55 | 0 |
| [mixpanel](../../packages/twin/mixpanel) | **11%** (14/123) | 10/10 | 4/109 | 0/4 | 0 |
| [cloudflare](../../packages/twin/cloudflare) | **3%** (85/3348) | 14/22 | 65/694 | 6/2632 | 0 |
| [sendgrid](../../packages/twin/sendgrid) | **1%** (3/403) | 3/3 | 0/392 | 0/8 | 0 |
| [notion](../../packages/twin/notion) | **1%** (1/179) | 1/45 | 0/78 | 0/56 | 132 |

*Across 105 twins: **8388/18651** verify-proven done (45%) · 354 claims out of date (a pack at protocol 1: its vendor half is unproven until it moves; generated/INDEX.md names each pack's protocol).*

<!-- capability-coverage:END -->

---

## The freshness cycle

> **Goal: stay close to vendor reality.** The whole sweep is a re-census campaign the owner
> dispatches; every merge carries the fast subset. Nothing here is scheduled — this repo has no
> CI and nothing is cron'd (see [AGENTS.md](../../AGENTS.md)).

```
        ┌─────────────────────────────────────────────────────────┐
 on ask │ 1. CAPTURE   real vendor responses → recorded-diff oracle │
        │ 2. CHECK     spec + recorded-diff + SDK + UI, per vendor   │
        │ 3. REPORT    one coverage report; deviations + gaps        │
        │ 4. GATE      fail the run on regressions; gaps → worklist  │
        └─────────────────────────────────────────────────────────┘
 per-merge:  the fast subset (spec + SDK + UI completeness) for the packs a change touched
```

1. **Capture** (refresh the oracle) — `bun scripts/capture/<v>.ts`. Read-only against
   the real vendor; rewrites the recorded-diff fixtures. A stale oracle is a false green, so a
   re-capture is part of any vendor re-census campaign. *(Today only `scripts/capture/slack.ts`
   exists.)*
2. **Check** — for each vendor, run every rung it holds over a seeded or freshly-pulled
   world. (Conformance checks the objects *present* in the world, so point it at real
   pulled state or a seed — an empty world checks nothing.)
3. **Report** — one coverage report: per-rung pass/fail, deviations, and the
   completeness "missing" set. This is the single artifact a human reads.
4. **Gate** — a regression (new violation, dropped coverage) fails the run; new gaps
   become the worklist. `scripts/twin-check.sh` is the full sweep, run on the owner's ask; a
   merge clears the touched packs' own suites plus the meta-gates it affects.

### Running it today

```bash
# per-vendor probe (one rung set), against a world at --root:
bun packages/twin/stripe/src/cli.ts conformance --root <world>

# the aggregate UI-rung runner (all 5 vendors, one GapReport) — wired into the gate below:
bun scripts/twin-conformance.ts --check-only

# the standing gate (all packs' conformance tests + tsc + cookbook + twin-conformance):
bash scripts/twin-check.sh
```

`scripts/twin-conformance.ts` (TWIN-99 / R-O6) is a real, wired-in gate step ("twin conformance
(ui)" in `scripts/twin-check.sh`): it sweeps the UI-completeness + UI-structure rungs for all 5
piloted vendors into one GapReport and fails the gate on any gap NOT already recorded, with a
reason, in the script's own `BASELINE` map — the known backlog (9 gaps today: 2 github, 1 jira, 6
slack) stays non-fatal, but a genuinely new regression turns the step red.

### What is still missing (open)

- **Capture scripts** for the four piloted vendors that lack one (github / stripe / jira /
  linear) — without them a re-census campaign has no oracle to refresh for those packs.
- **A single command** that runs capture → aggregate → report for a named vendor, so a
  re-census campaign is one dispatch per vendor rather than a hand-assembled sequence.

Note what is deliberately NOT missing: there is no schedule to build. Freshness here is an
owner-dispatched campaign, recorded as a campaign record — not a cron job nobody reads.

## Reference screenshots (UI proximity)

Reference screenshots are **not required for every screen.** The UI-conformance run
keeps a **searchable reference store** indexed by vendor + screen; on each run it
*looks one up* and, when a match exists, attaches it beside the mirror's own
screenshot for visual review. No match → you just get the mirror capture. References
are added opportunistically (whenever a real one is captured); search is what surfaces
them. Proximity is **never pixel-gated** — the gate is the structural checklist; the
screenshots are review artifacts.

## Known deviations

Every harness takes a **known-deviations allowlist**: a deviation we have inspected and
accept (with a written reason), so the gate stays green without hiding it. A deviation
must be either *fixed* or *explicitly allowlisted with a reason* — never silently
ignored. The allowlist is itself reviewable evidence of what the twin does **not**
faithfully reproduce.
