{"protocol":"clank-doc/1","frameworkVersion":"0.24.0","slug":"operational-recovery","title":"Recovery, controlled rollouts and operational evidence","description":"These features are opt in operator configuration. They add no runtime packages and make no calls to GitHub Actions. Local regression coverage is in operational foundations.test.mjs , point in time.test.mjs , data plane.test.mjs , and the ma","group":{"id":"security","title":"Security and resilience"},"url":"https://docs.clank.run/docs/operational-recovery","source":"docs/operational-recovery.md","headings":["Automated restore drills and operational alerts","Per commit point in time recovery","Managed canary deployments","Streaming usage and forecasts","Independently verifiable audit exports"],"tableOfContents":[{"id":"automated-restore-drills-and-operational-alerts","title":"Automated restore drills and operational alerts","level":2},{"id":"per-commit-point-in-time-recovery","title":"Per-commit point-in-time recovery","level":2},{"id":"managed-canary-deployments","title":"Managed canary deployments","level":2},{"id":"streaming-usage-and-forecasts","title":"Streaming usage and forecasts","level":2},{"id":"independently-verifiable-audit-exports","title":"Independently verifiable audit exports","level":2}],"markdown":"# Recovery, controlled rollouts and operational evidence\n\nThese features are opt-in operator configuration. They add no runtime packages and make no calls to GitHub Actions. Local regression coverage is in `operational-foundations.test.mjs`, `point-in-time.test.mjs`, `data-plane.test.mjs`, and the managed-canary cases in `platform.test.mjs`.\n\n## Automated restore drills and operational alerts\n\nPass `operations` to `openPlatform`:\n\n```ts\nconst platform = await openPlatform({\n  // Existing platform configuration...\n  operations: {\n    intervalMs: 60_000,\n    usageWarningPercent: 80,\n    notify: async (alert, signal) => {\n      // Deliver to your operator-selected destination. Deduplicate alert.id.\n      await incidentDestination.write(alert, { signal });\n    },\n    restoreDrills: {\n      intervalMs: 24 * 60 * 60_000,\n      timeoutMs: 30_000,\n      boot: async (projectId, { databasePath, signal }) => {\n        // Select trusted bootstrap code for this project. Open ONLY databasePath.\n        return bootDisposableApplication(projectId, databasePath, signal);\n      },\n      checks: [{ name: \"Application boot\", path: \"/healthz\", status: 200 }],\n    },\n  },\n});\n```\n\nA drill decrypts the newest retained backup, opens a private disposable SQLite copy, boots the supplied application, runs loopback HTTP checks and records its receipt. It never replaces production data. The callback must not load tenant JavaScript into a shared control-plane process, start production jobs, send real notifications or connect to production writable services. Use a trusted adapter or isolated runtime with those capabilities disabled. No drill can prove correctness of external data that was not backed up.\n\nThe monitor persists incident transitions and notification attempts for deployment failures, unhealthy always-on applications, overdue jobs, backup failures/overdue schedules, forecast usage warnings and failed drills. Notifications use leased at-least-once delivery, exponential backoff and stable IDs; a resolution cannot overtake its undelivered opening notification. Delivery callbacks must honor their abort signal. Operators can read `GET /api/admin/operations` or run a poll with `POST /api/admin/operations/run`; both require a platform-admin browser session and mutations require CSRF. Polling handles at most 1,000 projects/workspaces and five rotating job inspections per pass. Configure multiple operational domains if the platform exceeds those bounds. Setting `intervalMs: false` retains manual execution.\n\nDrill claims and results survive control-plane restarts. One due drill runs per pass. Failed setup retries after five minutes; completed reports preserve the rehearsal's phase, checks and timings. Independent encrypted object storage remains necessary to survive loss of the application volume.\n\n## Per-commit point-in-time recovery\n\n`openPointInTimeRecovery(database, options)` attaches a SQLite session to every write transaction. Encrypted changesets, their hash chain and the new state seal commit in the same SQLite transaction as application and service writes. A failed or oversized commit rolls back both. A consistent encrypted base backup and immutable exported journal files recover exact committed boundaries; this is not periodic snapshot approximation.\n\n```ts\nconst database = await openSQLite(schema, { path: \"/srv/project/data/app.sqlite\" });\n// For a NEW database, finish all service/schema initialization before attaching.\n// For an EXISTING journal, attach before opening services or admitting writes.\nconst recovery = await openPointInTimeRecovery(database, {\n  directory: \"/srv/project/recovery/epoch-2026-10\",\n  encryptionKey: keyFromSecretStore, // exactly 32 bytes; never store beside the archive\n  exportIntervalMs: 1_000,\n  maxTransactionBytes: 4 * 1024 * 1024,\n  maxStateBytes: 32 * 1024 * 1024,\n});\n// Admit application requests only after the awaited call completes.\n// At shutdown: stop application writers, await recovery.close(), database.close().\n```\n\nA recovery-enabled database must use one Clank connection. Another connection's write fences further managed writes through SQLite `data_version`; reopening verifies the canonical persisted state seal to detect changes while the process was down. Reopening a journal-bearing database without reattaching recovery rejects managed writes. Raw SQL tools, multiple writer processes, migrations and arbitrary connection access are outside the supported write contract. Administrative tools must stop the application and follow epoch rotation below. Restore a known verified epoch if an unexpected writer invalidates the seal; do not overwrite the stored seal to dismiss the error.\n\nAll captured tables require a primary key with no null values. Virtual tables, generated/hidden columns, more than 200,000 rows per table, and state exceeding `maxStateBytes` are rejected. SQLite integer values outside JavaScript's safe integer decoding range are rejected by the native reader. The logical state seal is recomputed for each transaction, including internal service reads that acquire the capture transaction; benchmark this opt-in bounded-database mode before production adoption. It is intentionally not a low-cost WAL archive for large databases. Trigger and foreign-key cascade effects are captured once; replay suspends their automatic effects, applies the captured changes, restores trigger definitions and checks foreign keys plus the canonical committed state seal.\n\nThe journal remains durable in the live SQLite database until epoch rotation. `flush()` fsyncs immutable encrypted entries and an authenticated export checkpoint. `status().committedThrough` may exceed `exportedThrough` while export is pending. Copy the entire repository, including `epoch.json`, `head.json`, journal entries and `base/`, to independently retained storage; loss of the live database before export loses those unexported transactions. A current independently retained checkpoint is necessary to detect replacement of an entire repository with an older valid copy. Commit timestamps are recorded immediately before commit and strictly ordered; exact sequence selection is the strongest boundary.\n\nRestore only into a stopped destination:\n\n```ts\nconst restored = await restorePointInTime({\n  directory: \"/restore/epoch-2026-10\",\n  encryptionKey: keyFromSecretStore,\n  targetPath: \"/srv/recovered/data/app.sqlite\",\n  throughSequence: 125, // alternatively asOf: timestampMilliseconds\n  confirmation: \"restore point in time\",\n});\n```\n\nRestore authenticates the base and every exported entry through the authenticated checkpoint, rejects gaps, conflicting changes, tampering and unavailable requested sequences, and replays inside the bounded SQLite worker namespace. It publishes the stopped destination only after verification. The original destination survives failed verification. `asOf` selects the latest available recorded commit at or before the timestamp; requests before the base snapshot fail. The recovered database starts without the old journal metadata so it can begin a new epoch.\n\nTo rotate for a schema migration: stop all writers; flush and retain the old encrypted epoch; restore its final sequence into a new stopped database; apply migrations to that new database; open it and initialize services; attach a new recovery repository; verify a restore drill; switch to the new database. This also bounds retained journal storage. To roll back the feature, stop writers and restore the latest verified sequence into a database without recovery metadata, then deploy the prior configuration. Do not drop journal metadata from a live writer or delete its archive as part of disabling the feature.\n\n## Managed canary deployments\n\n`openPlatform({ canary: ... })` enables measured traffic stages for local code-only deployments with an existing running release and managed ingress:\n\n```ts\ncanary: {\n  stages: [\n    { trafficPercent: 10, durationMs: 60_000, minimumSamples: 100 },\n    { trafficPercent: 50, durationMs: 60_000, minimumSamples: 100 },\n    { trafficPercent: 100, durationMs: 60_000, minimumSamples: 100 },\n  ],\n  maximumErrorRate: 0.01,\n  maximumP95Ms: 500,\n}\n```\n\nThe candidate first passes its health endpoint, then receives each configured percentage of requests. Assignment uses a deterministic 100-request cycle; there is no user/session affinity. Samples count completed responses actually routed to that candidate, using body completion latency, final status and stream errors. Cancelled client requests do not satisfy sample requirements. Up to the newest 10,000 samples per stage are retained; delayed completions from an earlier stage are excluded. Each stage must meet its duration, sample count, error and p95 limits. Missing traffic fails safely rather than silently promoting. Health, process liveness, current deployment authority and the renewable project lock are rechecked throughout. Failure removes candidate routing and the existing deployment rollback path stops it and restores prior background workers. The prior application remains the routing fallback. Final promotion uses the existing activation transaction and drains the prior release.\n\nBoth web versions share the same database during a code-only canary. Application code and data writes must be backward compatible. Previous background workers are quiesced before candidate workers start. Pending migrations are rejected for an existing running canary baseline; perform them through an explicit maintenance deployment with canaries disabled. Provider generations are not supported by this local parallel-runtime implementation and configured provider canary deployments fail closed. A first deployment has no baseline and follows normal health-gated activation.\n\n`GET /api/projects/:id/canary` returns durable stage and failure reports to an authorized project reader. Interrupted control-plane runs are marked interrupted on startup and retain the prior authoritative release. Disabling the option removes canary routing; durable reports remain. No migration is required for existing installations; the additive report table is created when configured.\n\nManaged local runtimes run beneath a separate supervisor process, outside the imported application. Loss of the control-plane IPC connection stops the runtime's process group or removes its exact Docker container. A private, fsynced `runtime-guardians` fence remains until cleanup is verified, and startup refuses to admit replacement writers while an earlier fence is unresolved. Preserve these files after an uncertain Docker/host failure; verify the recorded runtime has stopped before operator recovery. The process runner remains a trusted-application development mode, not a security boundary against applications that can access the host or deliberately escape their process group. Local crash regressions are in `platform-crash.test.mjs`; `platform-docker-crash.test.mjs` additionally verifies real container removal and database write order on an explicitly disposable Docker host.\n\nBackground recovery and incoming application requests share the same project lock, so requests wait for an in-progress recovery instead of launching a second runtime. After an abrupt controller failure during a locked operation, recovery also waits for the previous controller's lease to expire, which can take up to 30 seconds. A controller that loses its lease during startup stops its candidate without starting additional workers, publishing the runtime, or overwriting the successor's release status.\n\n## Streaming usage and forecasts\n\nIngress records response bytes when chunks are handed to the downstream reader, including terminal completion, cancellation and stream errors. It never bills a declared `Content-Length` as delivered bytes. HEAD and bodyless status responses count zero. Metrics emit once at body termination, and drain leases are released at that same boundary. This measures application payload consumed from managed ingress, not TLS, TCP, compression or remote receipt. Long streams are metered when they terminate, in the request's recorded usage period. Process death before termination can lose the unfinished metric.\n\nWorkspace usage payloads include `forecast.requests` and `forecast.transferBytes`, with observation duration, average rate, projected period total and estimated exhaustion time. Less than one hour of data reports insufficient evidence, while actual warning/exhaustion thresholds still apply. Forecasts extrapolate observed average demand and cannot predict bursts or future traffic changes.\n\n## Independently verifiable audit exports\n\nConfigure `auditExport` with an Ed25519 private key, stable key ID and an operator-controlled destination callback. The exporter signs sequence-linked canonical event digests into a durable outbox before delivery. Acknowledged batches advance the checkpoint; retries reuse the identical signed entries. Destination storage must deduplicate by sequence and digest and preserve an independent checkpoint. Keep the private key outside the platform database and protect trusted public keys through a separate configuration channel.\n\n`verifyAuditExport(entries, publicKeys, previousCheckpoint?)` verifies signatures, order, continuity, event content and allowed envelope fields and returns the next checkpoint. Omission, reordering and mutation are rejected. A verifier without an independently retained checkpoint cannot detect removal of an entire valid suffix or rollback to an older valid export. Existing audit IDs must be contiguous; historical deletion causes an explicit export error instead of claiming a complete chain. Plan retention and key rotation with the destination operator. `GET /api/admin/audit-export` exposes delivery status to browser platform admins. Removing the option stops export without deleting the audit log/outbox, and restoring the same keys and destination resumes it. Keys from earlier export periods must remain available to independent verifiers.\n"}