Files
parking_solution/wiki/concepts/backup-recovery.md
T
julian 9c6741a485
CI / check (push) Successful in 40s
docs(wiki): record backup deploy gotchas (compose allowlist + host mount)
Two lessons from the first park-buzi staging deploy, both in backup-recovery.md:
- A new server env var (BACKUP_KEY) must be added to docker-compose.yml's
  server.environment: allowlist, not just the Komodo secret/Stack env — otherwise
  the container never receives it (inspect shows it absent, not empty).
- The backup target must be a host path bind-mounted into the container; a desktop-
  automounted USB (/run/media/...) is invisible inside the container, so Test target
  reports 'does not exist'. Destinations are admin-provisioned (fstab + compose bind-
  mount), not operator-pluggable — partly a threat-model feature. Acknowledged as a
  flexibility limitation; USB-automount-to-container flow deferred.

Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
2026-06-29 17:38:24 +02:00

15 KiB

type, tags, sources, updated
type tags sources updated
concept
parking
durability
backup
recovery
security
crypto
2026-06-29

Backup & Disaster Recovery

The appliance's sqlite DB is the signed append-only-event-chain — the whole revenue/audit history. A disk failure or a stolen/destroyed PC currently means total loss (this is open-questions #5). This page is the settled design for an on-site, admin-driven backup that survives total hardware loss and restores to a fresh appliance with the signed chain still verifying. (Designed 2026-06-29.)

The recovery scenario it must satisfy

The driving scenario (the one that forces every decision below): the PC is gone — stolen or destroyed. Its SSD is LUKS-encrypted and TPM-sealed, so the disk is unrecoverable by design (a stolen disk won't unlock off its own TPM — see disk-os-hardening, tpm). We do not want the dead disk; we want to stand up a new PC, restore the backup, and continue signing the same chain. For that to work, recovery must depend on (a) the backup file and (b) two keys held out-of-band — never on the dead machine.

Key custody — the load-bearing decision

This is the part the whole plan rests on, and it interacts with the secure-element question (open-questions #6). Three independent keys, three custodians:

Key Lives Recoverable after PC loss? Job
EVENT_SIGNING_KEY fleet-deployment-komodo secret (park_buzi_event_signing_key), escrowed offsite Yes — by design Signs + verifies the ledger chain
park_buzi_backup_key (new) Komodo secret, escrowed offsite, separate from the signing key Yes Encrypts/decrypts the backup file
LUKS / TPM disk key The appliance's TPM only No — deliberately At-rest protection of the powered-off SSD
  • The signing key is decoupled from the TPM — kept an extractable software HMAC secret (append-only-event-chain, signer.ts), held in Komodo and escrowed by the operator. This is a conscious trade: a truly non-extractable TPM-sealed signing key (the #6 upgrade) would make the ledger unforgeable even against a host-root attacker — but it would also make the old ledger permanently unverifiable after total hardware loss (the sealed key dies with the machine; buildVerifier(keyId) would return undefined forever). You cannot have both "key can never be extracted" and "I can rescue the key after the machine dies" — they are the same property from two sides. Against the threat-model (the booth operator, who has a UI login, not host root) an escrowed software key is already tamper-evident, so the recoverable design is chosen today; revisiting #6 means re-accepting the unverifiable-after-loss cost. See tpm "TPM vs. ATECC608", fleet-deployment-komodo (the "EVENT_SIGNING_KEY-in-Core is a fraud-root blast radius" caveat is the same trade).

  • Backup key is separate from the signing key even though Komodo holds both — so they can be managed independently. Rationale: (1) the signing key must almost never rotate (every rotation fractures the chain into a new keyId segment — old events stay pinned to the old key forever), whereas the backup key may want routine rotation (a USB went home, a target was decommissioned); coupling them drags the cheap op into the expensive one. (2) The backup key travels to every backup destination (USB, NAS, SFTP); the signing key should travel nowhere but Komodo → process memory — sharing one key means every backup target conceptually exposes the signing key. (3) Keeping them separate keeps the #6 TPM-migration door open without re-wiring backups. Decided 2026-06-29 (the "one fewer secret to escrow" simplicity of a shared key is real, but weakest here because Komodo already holds both).

The keys are never inside the backup they unlock. A key can't decrypt the file it's locked in. Recovery = backup file + both escrowed keys, supplied out-of-band. The runbook must say this plainly so nobody "helpfully" stores the keys next to the backups.

What a backup contains

Full SQLite DB, snapshots included — one self-contained, restore-to-identical-appliance file (ledger + sessions + config + subscriptions + the entry-exit-points BLOBs). Chosen for completeness over size.

Size caveat (interacts with open-questions #10). Snapshot BLOBs dominate DB size and bloat every backup. They are unsigned, advisory, and already disk-pressure-pruned (entry-exit-points). A future "exclude snapshots" toggle (ledger/sessions/config only — much smaller, signed chain still fully preserved) is the obvious knob if backup size becomes a problem; the default is the complete picture.

The backup is produced via SQLite online-backup / VACUUM INTO (a consistent snapshot of the live WAL-mode DB — never a raw file copy, which can capture a torn WAL), then encrypted with park_buzi_backup_key. Acceptance test: a restored copy must still pass verifyChain — the signed chain is the thing being protected, so an unverifiable restore is a failed backup.

Triggers

  • Manual — an admin-only "Back up now" button runs immediately to the configured target.
  • Periodic — an in-process daily timer (same pattern as the snapshot-retention prune, entry-exit-points / snapshot-retention.ts): runs only if the configured target is reachable/mounted; surfaces last-success / last-error in the UI. No OS cron — it lives inside the Fastify process, works inside the container-deployment, and is configured in one place. (offline-first: the periodic path must tolerate a missing/unmounted target without failing the app.)

Destinations (admin-configurable)

All three supported in the first cut; the manual button and the periodic timer share them:

  • Local / USB / SATA disk — a mounted path on an attached disk. Simplest, fully offline, matches the air-gapped appliance. The strong first target.
  • Network drive (SMB/NFS) — a mounted share on the isolated LAN (a site NAS). Still local-network, no internet (network-isolation).
  • SFTP — push to an SFTP endpoint, useful for an offsite copy. FTP is excluded (plaintext credentials + data); SFTP is the safe equivalent.

The target must be a bind-mounted host path — NOT a casually-plugged USB (2026-06-29)

The server runs inside the parking-server container, so it can only stat()/write paths that are bind-mounted into that container. A USB stick the operator plugs in lands at a desktop auto-mount path on the host (/run/media/<user>/<UUID>), which does not exist inside the container — so the in-UI Test target correctly reports "location does not exist." This bit on the first booth deploy (2026-06-29): BACKUP_KEY was finally injected, then the target test failed because the USB path wasn't visible to the process.

So a backup destination is provisioned by the ADMIN at the host level, not chosen ad hoc by the operator. The procedure:

  1. Attach the disk (external HDD/SSD/USB) and mount it at a stable host path (e.g. /mnt/backup) via /etc/fstab by UUID — not the desktop automounter, whose UUID-named path changes per drive and vanishes on unplug.
  2. Bind-mount that host path into the container in the prod compose (e.g. /mnt/backup:/mnt/backup on the server service — same pattern as the /dev/usb printer passthrough in container-deployment).
  3. In the UI (Setup → Backup), set the target directory to the in-container path (/mnt/backup) and Test target — now writable.

This is partly a feature, not just a limitation (threat-model): because the destination is a host-provisioned bind-mount, the booth operator cannot redirect backups to a removable stick they walk off with — real destinations are an admin/host decision, on the trusted side of the trust-boundary. A network share (SMB/NFS) is the same shape: mount on the host, bind-mount in.

Limitation acknowledged: the backup target is therefore not operator-flexible — you cannot just plug in a USB and back up from the UI. Adding a new destination = a host fstab + compose bind-mount change + redeploy. For the appliance model (single-purpose, admin-provisioned) this is the right trade; a future "back up to a freshly-plugged removable drive" flow would need host-level automount detection wired to the container, which is deferred / not built.

Retention at the destination

Keep last N + thinned dailies (e.g. last 7 daily / last 4 weekly) — bounded disk use, and it survives the "a bad/partial run clobbered the only good copy" failure. (A single rolling overwrite-latest file was rejected for exactly that reason.)

Threat-model fit — restore is the dangerous half

Writing a backup is benign; restore is operator-adversary surface (threat-model). A restored DB replaces the live signed chain — so a malicious restore is a way to swap in a doctored history. Therefore:

  • Restore is NOT a booth button. It is an admin-only, out-of-band runbook action (new appliance, deliberate provisioning step), not something reachable from the operator console.
  • The backup target configuration and the "Back up now" action are admin-gated.
  • Backups do not weaken the chain's tamper-evidence: a restored chain is re-verified with the escrowed EVENT_SIGNING_KEY; a tampered restore fails verifyChain just as a tampered live DB would. The backup is a durability control, not an integrity one — integrity stays with the signed chain + reconciliation.

As-built (2026-06-29) — engine + local/mounted target

The first slice is built and tested: the backup engine + a local/mounted target + the daily timer + the manual route. What landed:

  • apps/server/src/backup.ts — the engine. Consistent online copy via better-sqlite3's native .backup() (a transactionally-consistent snapshot of the live WAL DB — not a raw file copy), then AES-256-GCM encryption with a scrypt-derived key from BACKUP_KEY. Self-describing header (magic | version | salt | iv | … | authTag) so a restore tool needs only the key + the file — zero new dependencies (Node crypto). The plaintext intermediate is written to scratch (not the removable/network target) and wiped in a finally, success or fail. Retention = keep-last-N + one-per-day-within-N-days (pruneOldBackups). Tested: round-trip decrypts to a byte-identical, queryable DB; a flipped byte or wrong key fails GCM auth; short key rejected; scratch plaintext always removed.
  • backup-service.ts — the target directory AND retention are admin-chosen in the UI (site_config.backup_target_dir, migration 0016; backup_keep_last + backup_keep_daily_days, migration 0017) and read fresh each run, so changing them takes effect with no restart. Retention columns are nullable → fall back to the code default (keep-last 7, keep-daily 30) per field. The encryption key is the ONLY backup env/Komodo secret (BACKUP_KEY) — a key must never live in the DB it backs up; target+retention are operational policy, not secrets. The service serializes concurrent runs (single in-flight guard) and records last-success / last-error; status() exposes targetDir, keepLast, keepDailyDays + keyPresent so the UI distinguishes "no target" from "no key".
  • routes/backup.ts — GET /api/backup/status (backup:read); PUT /api/backup/config to set/ clear the target (backup:update); POST /api/backup/test to probe a candidate path server-side — exists / is-a-dir / writable (backup:update); POST /api/backup/run (backup:create), a clean 409 backup_not_configured when target+key aren't both set. New backup permission resource (backup:read/update/create) in @parking/shared. No restore route — out-of-band by design.
  • apps/web/src/BackupSettings.tsx — a Setup → Backup tab (gated backup:read): an editable target-path field with a Test target probe (localized ok/missing/not-a-dir/not-writable), retention fields (keep-last / keep-daily-days), one Save, the status panel (config state, last-run size/pruned/error, a distinct amber missing BACKUP_KEY warning), a Back up now button, and the restore-is-out-of-band note. Full i18n (sq + en).
  • Komodo wiring. BACKUP_KEY is a per-booth Komodo secret ([[park_buzi_backup_key]] in komodo/resources.toml; documented in komodo/.env.komodo.example), escrowed offsite alongside EVENT_SIGNING_KEY. It is the only backup env var — target + retention are in the DB.

Gotcha — compose environment: is an ALLOWLIST (cost a full booth-deploy session, 2026-06-29). Wiring BACKUP_KEY as a Komodo secret + Stack-env line is necessary but not sufficient: docker-compose.yml's server.environment: block only forwards the variables it names. The key was wired everywhere (secret store, Stack env, .env.example, schema) but never added to that compose block, so the container came up without it — docker inspect ...Config.Env showed JWT_SECRET/EVENT_SIGNING_KEY present and BACKUP_KEY absent (not empty), while the Backup screen correctly reported "BACKUP_KEY missing". Diagnosis was muddied by chasing Komodo (secret name, re-sync, destroy/redeploy, env-only-change-doesn't-recreate) before checking the compose allowlist. Lesson: a new server env var needs a line in docker-compose.yml server.environment: too — that's the only place env reaches the container. Fixed: BACKUP_KEY: ${BACKUP_KEY:-} next to EVENT_SIGNING_KEY. Quick check on a booth: docker inspect <server> --format '{{range .Config.Env}}{{println .}}{{end}}' | grep -i backup.

  • server.ts — an unref'd daily timer (backupService.runScheduled), a no-op until configured, and deliberately NOT run at startup (a just-power-cut booth shouldn't write to a possibly-unmounted disk; the daily cadence + the manual button cover it).
  • Env documented in apps/server/.env.example (with the escrow + separate-key notes).

SMB/NFS already work — they're just a mounted path the admin enters as the target. Deferred to follow-up slices: an SFTP target and a restore runbook / CLI.

Status

Design settled 2026-06-29; engine + admin-configured local/mounted target + admin UI BUILT 2026-06-29 (SFTP + restore tooling pending). The target directory is admin-chosen in the UI (site_config, migration 0016), not an env var — the on-site admin picks where backups land; only BACKUP_KEY stays a server secret. Resolves the design half of open-questions #5 and the first build slices; records the key-custody stance that bears on #6 (signing stays decoupled from the TPM) and #10 (snapshots bloat backups → future exclude toggle). See append-only-event-chain, disk-os-hardening, tpm, fleet-deployment-komodo, reconciliation.