{"protocol":"clank-doc/1","frameworkVersion":"0.24.0","slug":"local-provider-fleet","title":"Local provider fleet simulator","description":"Use @clank.run/framework/fleet simulator to reproduce a fixed failure against an actual local coordinator, two Docker providers and two deployment agents. Each runs in its own OS process; transport uses authenticated loopback HTTP. This opt","group":{"id":"security","title":"Security and resilience"},"url":"https://docs.clank.run/docs/local-provider-fleet","source":"docs/local-provider-fleet.md","headings":["Run a scenario","Faults and evidence","CLI and reports","Cleanup and interruption"],"tableOfContents":[{"id":"run-a-scenario","title":"Run a scenario","level":2},{"id":"faults-and-evidence","title":"Faults and evidence","level":2},{"id":"cli-and-reports","title":"CLI and reports","level":2},{"id":"cleanup-and-interruption","title":"Cleanup and interruption","level":2}],"markdown":"# Local provider fleet simulator\n\nUse `@clank.run/framework/fleet-simulator` to reproduce a fixed failure against an actual\nlocal coordinator, two Docker providers and two deployment agents. Each runs in its own OS\nprocess; transport uses authenticated loopback HTTP. This optional module adds no dependency\nand is not imported by the browser entry point.\n\nRun drills only on an explicitly disposable Linux host with its own local Docker daemon and\ndedicated XFS project-quota mount. First obtain a current [Linux host certificate](linux-host-certification.md)\nfor the selected immutable image, non-root container identity, resources, quota and outbound\npolicy. The caller must run with both real and effective UID 0. Admission requires more than\nfive minutes of certificate validity; the final result also checks expiry. Changing the\nframework installation or selected policy invalidates the certificate.\n\n## Run a scenario\n\n```ts\nimport {\n  runLocalProviderFleetScenario,\n  exportLocalProviderFleetScenario,\n} from \"@clank.run/framework/fleet-simulator\";\n\nconst scenario = {\n  protocol: \"clank-fleet-scenario/1\",\n  kind: \"lease-loss\",\n} as const;\n\nconsole.log(exportLocalProviderFleetScenario(scenario));\nconst report = await runLocalProviderFleetScenario({\n  certificate: { directory: \"/operator/certification\", profile },\n  scenario,\n  disposable: true,\n  quotaIds: [1201, 1202], // Reserve two unused IDs on this dedicated mount.\n  portStart: 47000,       // Reserve ten unused application ports.\n  signal: controller.signal,\n});\nif (report.status !== \"passed\") throw new Error(report.reason);\n```\n\n`profile` is the exact certified `LinuxHostCertificationProfile`; `controller` is an\napplication-owned `AbortController`. Neither is part of the portable scenario. The two quota\nIDs must have zero existing byte/inode usage and limits. The API checks both before creating\nscratch data; it never clears a preexisting reservation. Ports are local to the disposable\nhost. Existing port listeners cause ordinary provider activation to fail.\n\n`parseLocalProviderFleetScenario(unknown)` validates plain static data and returns a frozen\nnormalized scenario. Unknown fields, getters, symbols, callbacks, unsupported faults and\nunbounded values are rejected. Defaults and inclusive integer bounds are:\n\n| Field | Default | Range |\n| --- | --- | --- |\n| `nodeTtlMs` | 5000 | 2000–30000 |\n| `operationLeaseMs` | 3000 | 1000–30000 |\n| `transportDelayMs` | 1500 | 200–10000 |\n\n`exportLocalProviderFleetScenario(scenario)` emits deterministic versioned JSON containing\nonly these fields and the fault kind. It excludes paths, image policy, credentials and reports.\nThe API captures operator inputs before awaiting host admission. It supports one bounded run\nat a time on a host and imposes a three-minute work deadline and bounded infrastructure calls.\n\n## Faults and evidence\n\n| Kind | Actual fault and required recovery |\n| --- | --- |\n| `takeover` | Kill the first agent with an outstanding claim, stop its owned provider, revoke its credentials, then let the second agent reclaim and execute with a higher fence. Old credentials must fail authentication, claim and observation. A separate stateful placement remains pinned to node A. |\n| `lease-loss` | Delay real lease-renewal HTTP responses beyond the operation lease. Wait for actual store expiry, reject stale completion and artifact access, then execute a reclaimed higher fence. |\n| `slow-transport` | Delay the real runtime-capsule HTTP response beyond the client deadline. Observe an actual transport timeout, restore transport and retry. |\n| `disk-read-only` | Remount the first provider's dedicated bind mount read-only inside its private mount namespace. Prove `EROFS`, reject deployment state writes, preserve the previously serving generation, remount writable and recover. |\n| `coordinator-restart` | Kill the coordinator and reopen its SQLite control store on the same listener. Preserve the prior desired generation and fence, then execute the next generation. |\n| `provider-restart` | Kill the provider while its Docker runtime exists. Reopen the same owned roots, remove its orphan runtime through production cleanup, then activate the next generation with its committed database. |\n\nEvery passing drill serves the fixed synthetic application through production ingress before\nand after recovery. It verifies the exact artifact digest, expected generation, synthetic row,\none committed migration, increasing fence and rejection of a stale provider request.\nCoordinator and provider restarts use persistent stores. The read-only fault affects the\nprovider's own namespace; it does not remount the Docker daemon's host filesystem.\n\nProvider workers run under the certified non-root application UID:GID, so storage ownership\nmatches the unmodified complete Docker provider service. The trusted provider processes retain\nonly `CAP_SYS_ADMIN`, `CAP_NET_ADMIN` and `CAP_DAC_OVERRIDE` for private bind mounts, quotas,\nowned nftables rules and local Docker access. They clear supplementary groups, restrict their\ncapability bounding set and set `NoNewPrivs`. Startup verifies those identities and capabilities.\nApplication containers retain the production launcher's `cap-drop=ALL` and non-root user.\nThe provider's access to the Docker daemon remains a privileged host-control boundary.\n`setpriv`, `mount` and `umount` are included in certificate host binding. See\n[Docker provider runtime](provider-docker-runtime.md) for ordinary provider identity setup.\n\nPortable takeover initializes the synthetic database on node B from the same immutable\nartifact. It proves portable placement and credential/operation fencing. The separate stopped\nstateful placement proves control-store pinning. These checks do not establish SQLite data\nreplication, promotion of a stateful database, or automatic supervisor leadership.\n\n## CLI and reports\n\nSave a mode-0600 JSON file owned by the invoking user with `certificate`, `scenario`, `quotaIds`\nand `portStart`, then run:\n\n```sh\nclank-provider fleet --config /operator/private-fleet.json --disposable\n```\n\nThe CLI accepts no arbitrary command or provider override. It rejects symlink configuration\nfiles, non-private permissions and files over 16 KiB. It prints the immutable\n`clank-fleet-report/1` result and exits 1 when blocked.\n\nReports contain seven checks (`host`, `processes`, `baseline`, `fault`, `recovery`, `fences`,\n`cleanup`), the verified artifact digest when available, and at most 512 sequenced events.\nEvents expose only elapsed time, fixed node/event names, generation and fence. No tokens,\nruntime capsules, database contents, operator paths or raw infrastructure errors are emitted.\nPossible blocked reasons are `certification-required`, `aborted`, `scenario-failed` and\n`cleanup-failed`. A partial recovery never counts as a passing scenario.\n\n## Cleanup and interruption\n\nNormal exit and cancellation stop agents before providers and the coordinator. Cleanup\nverifies the exact owned containers, ownership-checked network resources, scratch files and\nquota usage/limits. A passing report requires every check, including cleanup, to pass.\n\nThe root-owned `/run/clank-host-certification.attempt` marker excludes other fleet runs and\nhost-certification probes. Its private manifest records the process, boot, reserved quota IDs,\nsynthetic owner labels and, once created, the scratch root. Unconfirmed cleanup or a killed\nparent leaves this marker in place and blocks certificate admission. Children also shut down\non IPC disconnection. Never treat a missing parent as proof that cleanup finished.\n\nOn this disposable host, inspect the marker and verify its boot/process identity. Check only\nits recorded owners, network plans, scratch root and quota IDs using the production\nownership-checking helpers described in [Linux provider verification](linux-provider-isolation.md).\nRemove the marker only after proving those exact resources are gone. Re-certify after any\nhost/framework change. Destroying the disposable guest is an alternative to manual recovery.\nThere is no automatic marker unlock or broad Docker/network cleanup command.\n\nThe source repository's manual `tests/fixtures/fleet-simulator-guest.mjs` driver requires an\nexplicit disposable marker, fresh dedicated mount and preloaded immutable image. Ordinary\ntests exercise input, report and CLI admission; they do not substitute for privileged drills.\n"}