Volter World

What a twin is

A twin is a local, stateful replica of one vendor's API. Your real SDK talks to it exactly as it talks to the vendor, with a fake key, and gets the vendor's answer back. This page is the promise, in three parts, and the limits that come with it.

It speaks the vendor's API

A twin serves the vendor's wire: the same paths, the same request and response shapes, the same error bodies, the same pagination. It is checked against the vendor's own published spec, and against recordings of the real vendor, so that stripe, @octokit/rest, @slack/web-api, jira.js, @linear/sdk and the rest need no changes and no special mode.

What it does not model, it does not fake. A route the twin has not built fails the way the vendor would fail it, with the vendor's status and the vendor's error shape, never as a silent success. That is why a 404 or a 400 from a twin is information: either your app is wrong, or the twin is incomplete. The coverage page says, vendor by vendor, what fraction of the vendor's spec the twin serves, and each twin's README says which operations it does not serve yet.

It is stateful

A write you make is visible on the next read. A customer you create can be charged, listed, updated and deleted, with the vendor's rules about what is allowed when. Reads reflect writes because the twin keeps every write in a log and serves the log projected over its data; nothing is canned, and a scenario is never a fixed list of responses.

The state is yours to control. A fresh world starts from the twin's default data; your app's writes land on top of it; reset returns to the defaults; a branch records its own changes over the same history. Seed and reset and the model say how.

It is deterministic

A twin's answer is a function of the request and its state, byte for byte. Replay the same requests against the same starting state and you get the same responses, the same ids, the same timestamps, because time in a world is a clock the world sets, not the wall.

This is a promise about what a twin will never do. A simulated twin never reaches its vendor at serve time. A real-system root performs writes through the kernel and records receipts, as the model describes. Simulation never runs a language model. A generative vendor such as OpenAI or Anthropic serves a clearly labeled deterministic stub, or a scenario you script, so a test can prove that the call happened, that the tool calls and usage were shaped correctly, that state persisted and that the failure path works. It cannot prove anything about the quality of a real model's answer, and a twin that tried would stop being a test you can rerun.

What that means for you

  • Trust the wire, check the coverage. If your app uses an operation, look for it in the twin's census before you rely on the twin for it.
  • Read refusals as data. A refused route is either an app bug or a gap, never a twin pretending.
  • Assert on behavior, not on prose. For generative vendors, test the plumbing.
  • Expect the same answer twice. A differing result needs investigation of the test, twin state and runtime; determinism is a contract to verify, not proof that a twin cannot have a bug.