New concept page backup-recovery.md resolving the design half of open-question #5. Driving scenario: a stolen/destroyed PC whose LUKS+TPM disk is unrecoverable by design — recovery stands up a NEW PC, restores a backup, and keeps signing the SAME chain. Settled: admin-driven encrypted full-DB backup (SQLite online-backup/VACUUM INTO, snapshots included) to local/USB, SMB/NFS, or SFTP targets; manual button + an in-process daily timer; keep-last-N + dailies retention; restore is admin-only / out-of-band (operator-adversary surface). A restored copy must still verifyChain. Key custody (the load-bearing decision, bears on #6): three independent keys — EVENT_SIGNING_KEY kept an extractable, escrowed software key DECOUPLED from the TPM so the ledger survives total hardware loss (the conscious trade: a TPM-sealed signing key would be unforgeable but permanently unverifiable after the machine dies); a NEW dedicated park_buzi_backup_key in Komodo for backup encryption, separate from the signing key; the LUKS/TPM disk key, appliance-only and deliberately non-recoverable. Keys are never inside the backup they unlock. Updated open-questions #5 (design SETTLED) + #10 note; disk-os-hardening deploy runbook (why the signing key is not sealed + park_buzi_backup_key); index catalog + concept count. Design only — not yet built. Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
7.9 KiB
type, tags, sources, updated
| type | tags | sources | updated | ||||||
|---|---|---|---|---|---|---|---|---|---|
| concept |
|
2026-06-29 |
Backup & Disaster Recovery
The appliance's sqlite DB is the signed append-only-event-chain — the whole revenue/audit history. A disk failure or a stolen/destroyed PC currently means total loss (this is open-questions #5). This page is the settled design for an on-site, admin-driven backup that survives total hardware loss and restores to a fresh appliance with the signed chain still verifying. (Designed 2026-06-29.)
The recovery scenario it must satisfy
The driving scenario (the one that forces every decision below): the PC is gone — stolen or destroyed. Its SSD is LUKS-encrypted and TPM-sealed, so the disk is unrecoverable by design (a stolen disk won't unlock off its own TPM — see disk-os-hardening, tpm). We do not want the dead disk; we want to stand up a new PC, restore the backup, and continue signing the same chain. For that to work, recovery must depend on (a) the backup file and (b) two keys held out-of-band — never on the dead machine.
Key custody — the load-bearing decision
This is the part the whole plan rests on, and it interacts with the secure-element question (open-questions #6). Three independent keys, three custodians:
| Key | Lives | Recoverable after PC loss? | Job |
|---|---|---|---|
EVENT_SIGNING_KEY |
fleet-deployment-komodo secret (park_buzi_event_signing_key), escrowed offsite |
Yes — by design | Signs + verifies the ledger chain |
park_buzi_backup_key (new) |
Komodo secret, escrowed offsite, separate from the signing key | Yes | Encrypts/decrypts the backup file |
| LUKS / TPM disk key | The appliance's TPM only | No — deliberately | At-rest protection of the powered-off SSD |
-
The signing key is decoupled from the TPM — kept an extractable software HMAC secret (append-only-event-chain,
signer.ts), held in Komodo and escrowed by the operator. This is a conscious trade: a truly non-extractable TPM-sealed signing key (the #6 upgrade) would make the ledger unforgeable even against a host-root attacker — but it would also make the old ledger permanently unverifiable after total hardware loss (the sealed key dies with the machine;buildVerifier(keyId)would returnundefinedforever). You cannot have both "key can never be extracted" and "I can rescue the key after the machine dies" — they are the same property from two sides. Against the threat-model (the booth operator, who has a UI login, not host root) an escrowed software key is already tamper-evident, so the recoverable design is chosen today; revisiting #6 means re-accepting the unverifiable-after-loss cost. See tpm "TPM vs. ATECC608", fleet-deployment-komodo (the "EVENT_SIGNING_KEY-in-Core is a fraud-root blast radius" caveat is the same trade). -
Backup key is separate from the signing key even though Komodo holds both — so they can be managed independently. Rationale: (1) the signing key must almost never rotate (every rotation fractures the chain into a new
keyIdsegment — old events stay pinned to the old key forever), whereas the backup key may want routine rotation (a USB went home, a target was decommissioned); coupling them drags the cheap op into the expensive one. (2) The backup key travels to every backup destination (USB, NAS, SFTP); the signing key should travel nowhere but Komodo → process memory — sharing one key means every backup target conceptually exposes the signing key. (3) Keeping them separate keeps the #6 TPM-migration door open without re-wiring backups. Decided 2026-06-29 (the "one fewer secret to escrow" simplicity of a shared key is real, but weakest here because Komodo already holds both).
The keys are never inside the backup they unlock. A key can't decrypt the file it's locked in. Recovery = backup file + both escrowed keys, supplied out-of-band. The runbook must say this plainly so nobody "helpfully" stores the keys next to the backups.
What a backup contains
Full SQLite DB, snapshots included — one self-contained, restore-to-identical-appliance file (ledger + sessions + config + subscriptions + the entry-exit-points BLOBs). Chosen for completeness over size.
Size caveat (interacts with open-questions #10). Snapshot BLOBs dominate DB size and bloat every backup. They are unsigned, advisory, and already disk-pressure-pruned (entry-exit-points). A future "exclude snapshots" toggle (ledger/sessions/config only — much smaller, signed chain still fully preserved) is the obvious knob if backup size becomes a problem; the default is the complete picture.
The backup is produced via SQLite online-backup / VACUUM INTO (a consistent snapshot of the
live WAL-mode DB — never a raw file copy, which can capture a torn WAL), then encrypted with
park_buzi_backup_key. Acceptance test: a restored copy must still pass verifyChain — the
signed chain is the thing being protected, so an unverifiable restore is a failed backup.
Triggers
- Manual — an admin-only "Back up now" button runs immediately to the configured target.
- Periodic — an in-process daily timer (same pattern as the snapshot-retention prune,
entry-exit-points /
snapshot-retention.ts): runs only if the configured target is reachable/mounted; surfaces last-success / last-error in the UI. No OS cron — it lives inside the Fastify process, works inside the container-deployment, and is configured in one place. (offline-first: the periodic path must tolerate a missing/unmounted target without failing the app.)
Destinations (admin-configurable)
All three supported in the first cut; the manual button and the periodic timer share them:
- Local / USB / SATA disk — a mounted path on an attached disk. Simplest, fully offline, matches the air-gapped appliance. The strong first target.
- Network drive (SMB/NFS) — a mounted share on the isolated LAN (a site NAS). Still local-network, no internet (network-isolation).
- SFTP — push to an SFTP endpoint, useful for an offsite copy. FTP is excluded (plaintext credentials + data); SFTP is the safe equivalent.
Retention at the destination
Keep last N + thinned dailies (e.g. last 7 daily / last 4 weekly) — bounded disk use, and it survives the "a bad/partial run clobbered the only good copy" failure. (A single rolling overwrite-latest file was rejected for exactly that reason.)
Threat-model fit — restore is the dangerous half
Writing a backup is benign; restore is operator-adversary surface (threat-model). A restored DB replaces the live signed chain — so a malicious restore is a way to swap in a doctored history. Therefore:
- Restore is NOT a booth button. It is an admin-only, out-of-band runbook action (new appliance, deliberate provisioning step), not something reachable from the operator console.
- The backup target configuration and the "Back up now" action are admin-gated.
- Backups do not weaken the chain's tamper-evidence: a restored chain is re-verified with the
escrowed
EVENT_SIGNING_KEY; a tampered restore failsverifyChainjust as a tampered live DB would. The backup is a durability control, not an integrity one — integrity stays with the signed chain + reconciliation.
Status
Design settled 2026-06-29; not yet built. Resolves the design half of open-questions #5 (implementation pending), and records the key-custody stance that bears on #6 (signing stays decoupled from the TPM) and #10 (snapshots bloat backups → future exclude toggle). See append-only-event-chain, disk-os-hardening, tpm, fleet-deployment-komodo, reconciliation.