Test depth in the thin packs (TWIN-23 / architecture review §4.3, row C3)
The architecture review flagged eight packs as "thin" by a test-file-count proxy — test files ÷ source files: aws 14/35, livekit 2/14, anthropic 2/12, openai 2/11, vital 3/12, googlemaps/mapbox/openweather 2/10. This document is the honest, per-pack resolution the review's own §4.3 asked for: each pack either gained real test depth, or carries an explicit written acceptance of manifest-only coverage, with the specific grounds — never a bare "it's fine."
Why file-count is a poor proxy here (read before the per-pack sections)
A *-capabilities.ts manifest verify() is a test — it drives the twin's real request
handler through a create→read→assert-values round-trip, and (unlike a hand-authored unit test)
it is mutation-gated: scripts/mutation-test.ts swaps the pack's handler for a
behaviorally-dead saboteur and re-runs every done verify; a verify that still passes against
the dead twin is a survivor and fails the gate unless explicitly allowlisted with a named,
seam-independent reason (pure crypto, a static fixture, etc. — see the file's own header for the
full allowlist doctrine). That mutation gate runs as its own step in bash scripts/twin-check.sh
("mutation gate (teeth)") in --check mode, which is where this pass found it already wired in
— the July 2026 architecture review's P0 finding ("mutation harness unwired... appears nowhere in
twin-check.sh") predates that wiring and no longer describes the gate as it stands.
So for a pack with a thin *.test.ts file count but a rich manifest, "2 test files" understates
the real coverage: the manifest verifies ARE tests, they run in the gate, and they have
mutation-proven teeth against a dead twin. That is exactly what "manifest-only coverage" means in
this document, and it is a legitimate, provable form of coverage — enforced in the full gate's
--check mode — not a euphemism for untested.
That said, file-count-as-proxy still misses real gaps: helper modules that sit off the
handler's request path (CLI/service-bootstrap glue, connector pull/push edge cases, pure-function
input validation nothing ever calls with bad input) can be genuinely dark even when the manifest
is deep. This pass read each pack's manifest, its existing *.test.ts files, and its
non-test/non-capability source modules to find that residual real risk — see each section below.
How to read the per-pack sections
- Deepened — real behavioral tests were added (or, in two cases, a genuine bug the new test exposed was fixed). Each entry names the specific behavior now under test and why it was dark before.
- Accepted (manifest-only) — the specific grounds for why the manifest + mutation gate is sufficient for the surface in question, named explicitly (not "coverage is fine").
- Numbers in this document describe what this pass found and changed, not a live coverage
percentage — the manifests grow every cycle and a hardcoded "N done capabilities" figure would
rot immediately. For the current honest completeness number for any pack, run
bun scripts/twin-capabilities.tsor read the generated table in ./conformance.md.
aws (14 test files / 35 src files — least thin already)
Decision: accepted (manifest-only for the remaining gaps; already far past "thin" everywhere else). Grounds:
- aws is genuinely the deepest of the eight before this pass: 14 dedicated test files including
per-service
*-twin.test.ts(S3/DynamoDB/Timestream), per-service*-conformance.test.ts, per-service*-sigv4.test.ts(real HMAC-SigV4 crypto round-trips, hand-verified independent of the handler), and per-service*-sdk.integration.test.ts(the real AWS SDK client driving the twin unmodified). - All three sub-manifests (S3/DynamoDB/Timestream) are wired into
scripts/mutation-test.tswith their own handler seam (handleS3TwinRequest/handleDynamoDBTwinRequest/handleTimestreamTwinRequest) plus long, individually-reasoned allowlists for the verifies that legitimately bypass the handler (pure SigV4 math, pure XML builders/parsers, the DynamoDB expression engine called directly, the Timestream query engine called directly) — every one of those allowlist entries names the specific pure function it calls instead of the handler. - Complex engines that live in their own modules —
aws-dynamodb-expr.ts(condition/update expression evaluation),aws-dynamodb-update.ts,aws-timestream-query.ts,aws-s3-xml.ts— are each exercised BOTH transitively through the handler (via the twin/conformance tests) AND directly by dedicated capability verifies (the allowlist entries above), which is the two-path coverage this document elsewhere had to add by hand for the thinner packs. - The one narrow spot found:
aws.s3.select(S3 Select SQL-over-CSV,aws-s3-select.ts) has a singleniche-tier verify covering oneSELECT ... WHEREshape; the module supports more of the S3 Select SQL surface than that one verify exercises. Given the tier (niche), the size of the gap relative to aws's existing depth, and this pass's effort budget, this is called out here as a known, accepted narrow spot rather than deepened in this pass — a natural pickup for a future loop rotation if S3 Select gets real traffic.
No source changes in this pass.
livekit (2 test files / 14 src files)
Decision: mixed — deepened. The manifest (rooms/participants/egress/ingress/SIP/agent-dispatch,
mutation-gated via handleLiveKitTwinRequest, with a reasoned allowlist for the pure-crypto token
and webhook verifies) already covers the handler surface thoroughly, including many error/twirp
status-code branches. This pass found and closed the residual gap: the connector's own
confirm-guard logic, which sits below the handler and had no coverage at all.
- Added to
livekit-twin.test.ts(describe('connector', ...)):pushLiveKitAction's two "leave it pending" guards — a non-2xx real-vendor response, and a 2xx response whose echoed data doesn't actually match the requested change (e.g. egress reported as stillACTIVE, notCOMPLETE) — were previously untested; every existing push test's injected executor returned a matching success. A tamper probe (disabling the status-code guard) proved the new test goes red. - Added:
mintAccessToken's own input guards (missingapiKey/apiSecret;apiSecretshorter than the 32-byte minimum) —verifyAccessToken's failure modes were already thoroughly tested, the mint-side guards were not.
Accepted as manifest-only for the rest: the extensive twirpError validation/404/409/412 branches
across rooms/participants/egress/ingress/SIP, and the pure-crypto webhook/token round-trips, are
already mutation-gated and were confirmed already well covered.
anthropic (2 test files / 12 src files)
Decision: mixed — deepened, and a real bug found + fixed. anthropic-agents.ts (926 lines,
the Managed Agents control-plane: agents/environments/sessions/vaults/skills/deployments) has
no dedicated test file — its only prior coverage was the capability manifest's verify()
round-trips, which are real (mutation-gated, full CRUD+version+archive round-trips) but only
exercise the happy paths plus a few specific error cases per resource.
Added to anthropic-twin.test.ts (describe('managed agents — validation + lifecycle edge cases', ...)):
createAgentfield-size validation guards (name> 256 chars,tools.length> 128,skills.length> 20) and the object-formmodel: { id, speed }resolution — none of these was ever hit by theagents.crudverify or anything else.createSession'svault_idsvalidation (non-array → 400; unknown vault id → 404).createSession's own archived-agent 409 guard, called directly (previously the ONLY thing that ever exercised this class of check wasrunDeployment's separate duplicate pre-check, which short-circuits beforecreateSessionruns — socreateSession's own guard was dead as far as any verify/test was concerned).sendSessionEventsrejecting events sent to an archived session (409).
The last one caught a real bug, found by writing the test the manifest's own design comment
promised ("Sending to an archived session is a 409 conflict"): sendSessionEvents looked the
session up via getLive, which already excludes archived rows, so the archived-specific 409 check
immediately below it could never run — every archived-session request actually returned a generic
404, not the documented 409. Fixed in anthropic-agents.ts (sendSessionEvents now looks the
session up via getRow, which includes archived rows, so the 409 check is reachable); the fix is
two lines and doesn't touch any other code path. Confirmed via tamper probe (reverting to the old
getLive+404 behavior turns the new test red).
Accepted as manifest-only for the rest of the 926-line module: the full CRUD+version+archive round trips for agents/environments/sessions/vaults/skills, the webhook HMAC sign/verify+tamper, the multiagent thread-spawning, and the outcome-grading loop validation are all real, mutation-gated verifies and were confirmed already well covered by this pass's research.
openai (2 test files / 11 src files)
Decision: mixed — deepened, and a real bug found + fixed. openai-twin.test.ts is already
one of the deepest test files among the eight (streaming reconstruction, tool_calls/tool_choice
combinatorics, embeddings, moderations, logprobs, cursor pagination, batches, uploads, vector
stores, runs/threads/assistants), and the 100+-capability manifest is mutation-gated with a single,
narrowly-reasoned allowlist entry (webhook HMAC, pure local crypto). This pass's research found one
genuine correctness bug in the untested remainder:
- Fixed: the chat-completion
stopsequence truncation loop (openai-twin.ts,buildChoice) broke at whicheverstopstring was listed FIRST in the array that matched anywhere in the text, rather than at whichever stop string occurs EARLIEST in the generated text. The only existing capability (openai.chat.stop) only ever passes a single-elementstoparray, so this array-order bug was invisible to it. Fixed to track the minimum match index across the wholestoplist. Added a test (chat params + new families (handler)inopenai-twin.test.ts) with two stop strings in reverse-of-occurrence order, asserting truncation happens at the earlier one; confirmed via tamper probe (reverting to the old break-on-first-array-element logic turns it red).
Accepted as manifest-only / deferred (found, not fixed, in this pass — narrower and lower-impact
than the stop-sequence bug, a reasonable pickup for a future rotation): a fine-tuning pause
status guard that's never exercised against an already-cancelled job (the manifest only tests
pause→resume on a healthy job), and hyperparameters echo-vs-default branch on fine-tuning job
creation (no capability or test ever passes a custom hyperparameters payload).
vital (3 test files / 12 src files)
Decision: mixed — deepened. The 60+-capability manifest (mutation-gated via
handleVitalTwinRequest, with a reasoned allowlist for pure svix-HMAC webhook functions and the
zero-import buildPscAvailability) covers the lab-test/order/PSC/webhook surface well. This pass
closed two residual gaps in vital-data.ts/vital-twin.ts that no verify or test reached:
dateWindow(backs every summary endpoint: sleep/activity/workouts/body) has three fallback/clamp branches — an inverted range (end before start) falling back to a single day, an unparseable date string falling back to a single day rather than throwing, and a 31-day cap on wide ranges — that every existing capability/test left dark (all of them pass a valid, same-or-adjacent-day window). Added three tests exercising each branch through the real/v3/summary/*endpoints. Tamper probe (removing the 31-day cap) confirmed the wide-range test goes red.GET /v3/order/:id/results' populated-results branch (order.status === 'completed') was reachable only via a direct state seed — there is no write-API path in the twin that ever transitions an order tocompleted(the real Vital lifecycle completes asynchronously once the lab processes the sample), so every existing verify/test only ever saw the empty-results default. Added a test that seeds a completed order directly viaapplyTwinWrite(the documented test-controlseedidiom (the seed guidance) for reaching states the write API can't construct) and asserts the populated branch serves the seeded results.
One finding from this pass's research was deliberately not turned into a test:
vital.lab_tests.markers_pagination's verify only asserts the echoed page/size metadata
fields, while markersEnvelope (vital-catalog.ts) always returns the full, unsliced markers
array regardless of page/size — the pagination is metadata-only. This is a real fidelity gap
(the twin doesn't actually paginate the array), but changing markersEnvelope's behavior is a
product-behavior change outside this pass's test-depth scope, and asserting today's actual
(unsliced) behavior in a new test would just be pinning the gap rather than surfacing it. Recorded
here as an explicit, honest known-gap for a future pass rather than silently fixed or silently
tested-around.
googlemaps (2 test files / 10 src files)
Decision: mixed — deepened. The 90+-capability manifest is mutation-gated via
handleGoogleMapsTwinRequest with no allowlist exceptions (every done verify is expected to
route through the handler). This pass closed two auth/pagination edge cases in
googlemaps-auth.ts/googlemaps-twin.ts that no verify or test reached:
- The
request-deniedsentinel key (REQUEST_DENIEDwith a "not authorized" message) — distinct from a missing key or theOVER_QUERY_LIMIT/OVER_DAILY_LIMITsentinels, which the manifest does cover — was never sent by anything. Added a test. decodePageToken's failure path (a corrupted/garbagepagetoken=) —googlemaps.places.paginationonly round-trips a token the twin itself generated. Added a test sending a malformed token and assertingINVALID_REQUESTrather than a crash or a silently-wrong page.
Accepted as manifest-only for the rest: missing-key REQUEST_DENIED, the quota sentinels,
DST-offset math (both branches), haversine/elevation/plus-code happy paths, and the
photo_reference invalid-request path were all confirmed already well covered by the manifest.
mapbox (2 test files / 10 src files)
Decision: mixed — deepened. The 80+-capability manifest is mutation-gated via
handleMapboxTwinRequest with no allowlist exceptions. This pass closed two edge cases in
mapbox-auth.ts/mapbox-data.ts that no verify or test reached:
- The
forbiddensentinel access token (→ HTTP 403) — distinct from a missing/malformed token (401) or therate-limitedsentinel (429), both of which the manifest covers. Added a test. parseLngLat's numeric RANGE check (±90/±180) — every existing "invalid coordinate" capability sends either a missing param or a non-numeric string; none sends well-formed-but-out-of-range numbers, so the range clamp itself was reachable but never exercised. Added a test via the isochrone endpoint (/isochrone/v1/mapbox/driving/200,100), which callsparseLngLatdirectly and returns 422 on anullresult. Tamper probe (disabling the range clamp) confirmed the test goes red.
Accepted as manifest-only for the rest: missing/invalid-token 401s, isochrone contour validation, matrix coordinate caps, static-image size caps, and tilesets/tokens/uploads required-field validation were all confirmed already well covered.
openweather (2 test files / 10 src files)
Decision: mixed — deepened. The 50+-capability manifest is mutation-gated via
handleOpenWeatherTwinRequest with no allowlist exceptions. This pass closed one edge case in
openweather-twin.ts (handleAirHistory) that no verify or test reached:
- The
end <= startinverted-range rejection (400 "end must be greater than start") —openweather.air.historyonly ever sends a valid, correctly-ordered range, andopenweather.air.history_requires_rangeonly tests a MISSING start/end, never an inverted one. Added a test for the inverted-range rejection, plus a companion test for the adjacent 200-sample truncation cap on wide ranges (same function, same gap in coverage).
Accepted as manifest-only for the rest: parseLatLon's out-of-range check IS already exercised
(shared by every lat/lon endpoint via openweather.current.invalid_coords), and the
missing/invalid/rate-limited auth sentinels, zip-not-found 404, and onecall
exclude/day_summary/timemachine paths were all confirmed already well covered.
Summary
| Pack | Test files / src files | Decision | What changed this pass |
|---|---|---|---|
| aws | 14 / 35 | accepted (manifest-only) | none — already the deepest pack; one narrow niche gap (S3 Select) named, not closed |
| livekit | 2 / 14 | deepened | connector push-confirm guards + mintAccessToken input guards |
| anthropic | 2 / 12 | deepened + bug fix | Managed Agents validation/lifecycle edge cases; fixed a dead 409 guard in sendSessionEvents |
| openai | 2 / 11 | deepened + bug fix | fixed + tested multi-stop-sequence earliest-match ordering; two narrower gaps named, deferred |
| vital | 3 / 12 | deepened | dateWindow fallback/cap edge cases; seeded-state test for the populated-results branch; one fidelity gap (pagination theater) named, deferred |
| googlemaps | 2 / 10 | deepened | request-denied auth sentinel; malformed-pagetoken rejection |
| mapbox | 2 / 10 | deepened | forbidden auth sentinel; out-of-range coordinate rejection |
| openweather | 2 / 10 | deepened | air-pollution history inverted-range rejection + truncation cap |
Every pack above either gained real, behavior-asserting tests for a specific named gap, or carries specific, checkable grounds for why the manifest + mutation gate already covers the surface in question — satisfying TWIN-23 dev/01 for all eight enumerated packs.