Files
parking_solution/wiki/concepts/backup-recovery.md
T
julian 84f00db48b
Build desktop / desktop (push) Successful in 4m17s
Build & push images / images (push) Failing after 39s
CI / check (push) Successful in 39s
feat(backup): admin-tunable retention + BACKUP_KEY as a Komodo secret
Retention (keep-last / keep-daily-days) is operational policy the on-site admin
should tune, not a server env var requiring a redeploy -- same reasoning that moved
the target directory to the UI.

- Migration 0017: site_config.backup_keep_last + backup_keep_daily_days (nullable;
  null = code default 7 / 30 per field).
- BackupService reads retention fresh each run; status() exposes keepLast +
  keepDailyDays. DEFAULT_BACKUP_RETENTION is now a pure code default (env reads gone).
- PUT /api/backup/config accepts keepLast / keepDailyDays (non-negative int, or null
  to reset to default; 400 on negative).
- UI: two retention fields on the Backup config card; one Save covers target +
  retention. i18n sq + en.

BACKUP_KEY wired into Komodo:
- komodo/resources.toml: BACKUP_KEY=[[park_buzi_backup_key]] (per-booth secret,
  alongside JWT / signing keys).
- komodo/.env.komodo.example: documents it as the ONLY backup env var -- escrow it
  offsite alongside EVENT_SIGNING_KEY (recovery needs both); target + retention are
  admin-chosen in the UI / DB, not env. Server .env.example trimmed to just BACKUP_KEY.

Also carries the small in-progress setup-intro i18n copy trim.

Tests: 218 server tests green, incl. retention persist / reset-to-default / reject-
negative and the updated status shape. Migration applies cleanly (needed a
statement-breakpoint between the two ALTERs). Wiki backup-recovery updated.

Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
2026-06-29 12:52:18 +02:00

12 KiB

type, tags, sources, updated
type tags sources updated
concept
parking
durability
backup
recovery
security
crypto
2026-06-29

Backup & Disaster Recovery

The appliance's sqlite DB is the signed append-only-event-chain — the whole revenue/audit history. A disk failure or a stolen/destroyed PC currently means total loss (this is open-questions #5). This page is the settled design for an on-site, admin-driven backup that survives total hardware loss and restores to a fresh appliance with the signed chain still verifying. (Designed 2026-06-29.)

The recovery scenario it must satisfy

The driving scenario (the one that forces every decision below): the PC is gone — stolen or destroyed. Its SSD is LUKS-encrypted and TPM-sealed, so the disk is unrecoverable by design (a stolen disk won't unlock off its own TPM — see disk-os-hardening, tpm). We do not want the dead disk; we want to stand up a new PC, restore the backup, and continue signing the same chain. For that to work, recovery must depend on (a) the backup file and (b) two keys held out-of-band — never on the dead machine.

Key custody — the load-bearing decision

This is the part the whole plan rests on, and it interacts with the secure-element question (open-questions #6). Three independent keys, three custodians:

Key Lives Recoverable after PC loss? Job
EVENT_SIGNING_KEY fleet-deployment-komodo secret (park_buzi_event_signing_key), escrowed offsite Yes — by design Signs + verifies the ledger chain
park_buzi_backup_key (new) Komodo secret, escrowed offsite, separate from the signing key Yes Encrypts/decrypts the backup file
LUKS / TPM disk key The appliance's TPM only No — deliberately At-rest protection of the powered-off SSD
  • The signing key is decoupled from the TPM — kept an extractable software HMAC secret (append-only-event-chain, signer.ts), held in Komodo and escrowed by the operator. This is a conscious trade: a truly non-extractable TPM-sealed signing key (the #6 upgrade) would make the ledger unforgeable even against a host-root attacker — but it would also make the old ledger permanently unverifiable after total hardware loss (the sealed key dies with the machine; buildVerifier(keyId) would return undefined forever). You cannot have both "key can never be extracted" and "I can rescue the key after the machine dies" — they are the same property from two sides. Against the threat-model (the booth operator, who has a UI login, not host root) an escrowed software key is already tamper-evident, so the recoverable design is chosen today; revisiting #6 means re-accepting the unverifiable-after-loss cost. See tpm "TPM vs. ATECC608", fleet-deployment-komodo (the "EVENT_SIGNING_KEY-in-Core is a fraud-root blast radius" caveat is the same trade).

  • Backup key is separate from the signing key even though Komodo holds both — so they can be managed independently. Rationale: (1) the signing key must almost never rotate (every rotation fractures the chain into a new keyId segment — old events stay pinned to the old key forever), whereas the backup key may want routine rotation (a USB went home, a target was decommissioned); coupling them drags the cheap op into the expensive one. (2) The backup key travels to every backup destination (USB, NAS, SFTP); the signing key should travel nowhere but Komodo → process memory — sharing one key means every backup target conceptually exposes the signing key. (3) Keeping them separate keeps the #6 TPM-migration door open without re-wiring backups. Decided 2026-06-29 (the "one fewer secret to escrow" simplicity of a shared key is real, but weakest here because Komodo already holds both).

The keys are never inside the backup they unlock. A key can't decrypt the file it's locked in. Recovery = backup file + both escrowed keys, supplied out-of-band. The runbook must say this plainly so nobody "helpfully" stores the keys next to the backups.

What a backup contains

Full SQLite DB, snapshots included — one self-contained, restore-to-identical-appliance file (ledger + sessions + config + subscriptions + the entry-exit-points BLOBs). Chosen for completeness over size.

Size caveat (interacts with open-questions #10). Snapshot BLOBs dominate DB size and bloat every backup. They are unsigned, advisory, and already disk-pressure-pruned (entry-exit-points). A future "exclude snapshots" toggle (ledger/sessions/config only — much smaller, signed chain still fully preserved) is the obvious knob if backup size becomes a problem; the default is the complete picture.

The backup is produced via SQLite online-backup / VACUUM INTO (a consistent snapshot of the live WAL-mode DB — never a raw file copy, which can capture a torn WAL), then encrypted with park_buzi_backup_key. Acceptance test: a restored copy must still pass verifyChain — the signed chain is the thing being protected, so an unverifiable restore is a failed backup.

Triggers

  • Manual — an admin-only "Back up now" button runs immediately to the configured target.
  • Periodic — an in-process daily timer (same pattern as the snapshot-retention prune, entry-exit-points / snapshot-retention.ts): runs only if the configured target is reachable/mounted; surfaces last-success / last-error in the UI. No OS cron — it lives inside the Fastify process, works inside the container-deployment, and is configured in one place. (offline-first: the periodic path must tolerate a missing/unmounted target without failing the app.)

Destinations (admin-configurable)

All three supported in the first cut; the manual button and the periodic timer share them:

  • Local / USB / SATA disk — a mounted path on an attached disk. Simplest, fully offline, matches the air-gapped appliance. The strong first target.
  • Network drive (SMB/NFS) — a mounted share on the isolated LAN (a site NAS). Still local-network, no internet (network-isolation).
  • SFTP — push to an SFTP endpoint, useful for an offsite copy. FTP is excluded (plaintext credentials + data); SFTP is the safe equivalent.

Retention at the destination

Keep last N + thinned dailies (e.g. last 7 daily / last 4 weekly) — bounded disk use, and it survives the "a bad/partial run clobbered the only good copy" failure. (A single rolling overwrite-latest file was rejected for exactly that reason.)

Threat-model fit — restore is the dangerous half

Writing a backup is benign; restore is operator-adversary surface (threat-model). A restored DB replaces the live signed chain — so a malicious restore is a way to swap in a doctored history. Therefore:

  • Restore is NOT a booth button. It is an admin-only, out-of-band runbook action (new appliance, deliberate provisioning step), not something reachable from the operator console.
  • The backup target configuration and the "Back up now" action are admin-gated.
  • Backups do not weaken the chain's tamper-evidence: a restored chain is re-verified with the escrowed EVENT_SIGNING_KEY; a tampered restore fails verifyChain just as a tampered live DB would. The backup is a durability control, not an integrity one — integrity stays with the signed chain + reconciliation.

As-built (2026-06-29) — engine + local/mounted target

The first slice is built and tested: the backup engine + a local/mounted target + the daily timer + the manual route. What landed:

  • apps/server/src/backup.ts — the engine. Consistent online copy via better-sqlite3's native .backup() (a transactionally-consistent snapshot of the live WAL DB — not a raw file copy), then AES-256-GCM encryption with a scrypt-derived key from BACKUP_KEY. Self-describing header (magic | version | salt | iv | … | authTag) so a restore tool needs only the key + the file — zero new dependencies (Node crypto). The plaintext intermediate is written to scratch (not the removable/network target) and wiped in a finally, success or fail. Retention = keep-last-N + one-per-day-within-N-days (pruneOldBackups). Tested: round-trip decrypts to a byte-identical, queryable DB; a flipped byte or wrong key fails GCM auth; short key rejected; scratch plaintext always removed.
  • backup-service.ts — the target directory AND retention are admin-chosen in the UI (site_config.backup_target_dir, migration 0016; backup_keep_last + backup_keep_daily_days, migration 0017) and read fresh each run, so changing them takes effect with no restart. Retention columns are nullable → fall back to the code default (keep-last 7, keep-daily 30) per field. The encryption key is the ONLY backup env/Komodo secret (BACKUP_KEY) — a key must never live in the DB it backs up; target+retention are operational policy, not secrets. The service serializes concurrent runs (single in-flight guard) and records last-success / last-error; status() exposes targetDir, keepLast, keepDailyDays + keyPresent so the UI distinguishes "no target" from "no key".
  • routes/backup.ts — GET /api/backup/status (backup:read); PUT /api/backup/config to set/ clear the target (backup:update); POST /api/backup/test to probe a candidate path server-side — exists / is-a-dir / writable (backup:update); POST /api/backup/run (backup:create), a clean 409 backup_not_configured when target+key aren't both set. New backup permission resource (backup:read/update/create) in @parking/shared. No restore route — out-of-band by design.
  • apps/web/src/BackupSettings.tsx — a Setup → Backup tab (gated backup:read): an editable target-path field with a Test target probe (localized ok/missing/not-a-dir/not-writable), retention fields (keep-last / keep-daily-days), one Save, the status panel (config state, last-run size/pruned/error, a distinct amber missing BACKUP_KEY warning), a Back up now button, and the restore-is-out-of-band note. Full i18n (sq + en).
  • Komodo wiring. BACKUP_KEY is a per-booth Komodo secret ([[park_buzi_backup_key]] in komodo/resources.toml; documented in komodo/.env.komodo.example), escrowed offsite alongside EVENT_SIGNING_KEY. It is the only backup env var — target + retention are in the DB.
  • server.ts — an unref'd daily timer (backupService.runScheduled), a no-op until configured, and deliberately NOT run at startup (a just-power-cut booth shouldn't write to a possibly-unmounted disk; the daily cadence + the manual button cover it).
  • Env documented in apps/server/.env.example (with the escrow + separate-key notes).

SMB/NFS already work — they're just a mounted path the admin enters as the target. Deferred to follow-up slices: an SFTP target and a restore runbook / CLI.

Status

Design settled 2026-06-29; engine + admin-configured local/mounted target + admin UI BUILT 2026-06-29 (SFTP + restore tooling pending). The target directory is admin-chosen in the UI (site_config, migration 0016), not an env var — the on-site admin picks where backups land; only BACKUP_KEY stays a server secret. Resolves the design half of open-questions #5 and the first build slices; records the key-custody stance that bears on #6 (signing stays decoupled from the TPM) and #10 (snapshots bloat backups → future exclude toggle). See append-only-event-chain, disk-os-hardening, tpm, fleet-deployment-komodo, reconciliation.