Files
parking_solution/wiki/decisions/appliance-provisioning.md
T
julian 7eadf71a0b docs(wiki): appliance-provisioning — Komodo deploy is now the primary flow
§6 split: §6 = Docker engine only; new §7 = the Komodo Periphery deploy (PRIMARY,
verified end-to-end on park-buzi 2026-06-27):
- 7a install Periphery (onboarding key, user-mode/outbound, runs as admin, no
  inbound port; core_address = Core's proxy URL)
- 7b deploy the Stack in Core (registry+git accounts, per-booth [[..]] secrets,
  env incl. COOKIE_SECURE=0; seed admin via Komodo's container terminal — no SSH)
- 7b-bis fleet-as-code via komodo/resources.toml + ResourceSync (empty diff =
  in sync)
- 7c break-glass: manual booth.sh when mesh/Core is down

Added Komodo deploy gotchas 7-11 (core_address is the proxy URL not :9120;
git-auth ≠ registry-auth; user-mode vs /etc/komodo root_directory; core_address
singular; empty-diff/disabled-Execute = success). §5b SSH TODO reframed (Komodo
removes SSH from routine ops). Header + date updated; log entry added. Fixed a
stale [[atecc608-secure-element]] alias in the prior log entry.

Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
2026-06-27 12:30:08 +02:00

19 KiB
Raw Blame History

type, tags, sources, updated, status
type tags sources updated status
reference
parking
deployment
appliance
hardening
runbook
offline-first
2026-06-27 settled

Appliance provisioning runbook (booth PC)

Step-by-step to take a booth PC from factory Windows to a hardened, encrypted, container-running parking appliance. Written from the first real provisioning, 2026-06-23 (hardening) + first Komodo deploy, 2026-06-27 (the runtime) — every command here was run and verified on the actual hardware, including the firmware-specific workaround. Companion to disk-os-hardening (the why), tpm (TPM analysis), container-deployment (the images), and fleet-deployment-komodo (the deploy control plane this runbook's §7 uses).

⚠ This box is the threat-model defence. The load-bearing anti-fraud control is still reconciliation over the append-only-event-chain — disk encryption + Secure Boot raise the cost of offline tamper, they don't replace reconciliation.

Reference hardware (first unit, 2026-06-23)

  • Dell OptiPlex 7070, Intel Core i5-8500 (Coffee Lake), 238 GB SATA SSD (/dev/sda).
  • TPM 2.0 — discrete Nuvoton (Get-Tpm → ManufacturerIdTxt NTC, fw 7.2.1.0). NOT Intel PTT/fTPM. Discrete ⇒ an external LPC/SPI bus exists (bus-sniff is a theoretical physical attack on PCR-only sealing — accepted; see tpm). Used PC — previous owner irrelevant.
  • Shipped Windows 11; formatted to Ubuntu 26.04 LTS (the decided platform — desktop-shell-tauri).
  • TPM: leave On. Used PC → Clear the TPM once (Security → TPM → Clear) so the prior owner's keys are wiped before LUKS enrollment. (PPI "Bypass for Clear" was unchecked → it asks for physical confirmation at next boot; that's normal.)
  • Secure Boot: Enabled, Deployed Mode (not Audit). "Enable Custom Mode" UNCHECKED = Standard Mode with stock Microsoft keys — this is what Ubuntu's signed shim needs. Do NOT touch PK/KEK/db/dbx. NB: the 7070's Expert Key Management is edit-only (Save/Replace/Append/Delete — no read-only "View Key"), so you cannot inspect db from BIOS; verify via the live USB instead (step 2).
  • Boot: UEFI only (no CSM/Legacy — a Legacy install has no Secure Boot / TPM-seal path).
  • Set a BIOS admin password.

2. Boot the Ubuntu 26.04 USB (Secure Boot ON)

  • Flash the ISO DIRECTLY (Rufus GPT/UEFI, Etcher, or dd). NOT Ventoy — Ventoy's own bootloader isn't in db, so Secure Boot rejects it with Verification failed: (0x1A) Security Violation (this is Secure Boot working correctly, not a fault). A directly-flashed Ubuntu USB boots the Microsoft-signed shim, which stock db trusts.
  • F12 at the Dell logo → pick the USB under UEFI BOOT.
  • Reaching the installer with Secure Boot ON = positive proof the MS third-party UEFI CA is in db (the verification the BIOS couldn't show us).

3. Encrypted install — the firmware workaround (IMPORTANT)

The 26.04 installer disk page offers: No Encryption / Encrypt with a passphrase / Use hardware-backed encryption (+ advanced LVM/ZFS, both ZFS experimental).

  • "Use hardware-backed encryption" FAILS on this 7070 with: PCR_UNUSABLE … error with secure boot policy (PCR7) measurements: generating secure boot profiles for systems with timestamp revocation (dbt) support is currently not supported. → Ubuntu's automated FDE profiler can't model PCR7 on Dell firmware carrying a dbt (UEFI timestamp revocation list). It is NOT a TPM or Secure-Boot fault — both are fine.
  • So: choose "Encrypt with a passphrase". Set a strong passphrase and SAVE IT OFF-MACHINE (phone / password manager). It is both the boot unlock (until TPM sealing) AND the permanent recovery slot. Finish the install.
  • Result (verify with lsblk): sda1 vfat /boot/efi, sda2 ext4 /boot, sda3 crypto_LUKS → dm_crypt-0 (LVM2) → ubuntu--vg-ubuntu--lv ext4 /.

4. Seal LUKS to the TPM (manual — PCR 7 only)

Do this AFTER first boot. Manual enrollment sidesteps the installer's dbt profiler and lets us pick PCRs. Bind to PCR 7 only (Secure Boot state): it catches the attack that matters (disabling Secure Boot to boot a tampered kernel) WITHOUT breaking on routine kernel/GRUB updates (which churn PCRs 4/8/9 → would otherwise drop every boot to the passphrase). Firmware-only PCR 0 is the fallback if PCR 7 ever errors.

sudo apt update && sudo apt install -y tpm2-tools
sudo tpm2_pcrread sha256                 # sanity: PCRs 0-10 populated, PCR 7 has a real value

# Enroll the TPM (prompts for the EXISTING install passphrase to authorize the new slot):
sudo systemd-cryptenroll --tpm2-device=auto --tpm2-pcrs=7 /dev/sda3

# Verify TWO slots — keep BOTH (slot 0 password = recovery, slot 1 tpm2 = auto-unlock):
sudo systemd-cryptenroll /dev/sda3
#   SLOT TYPE
#      0 password
#      1 tpm2

Wire it into boot (back up first; the mapping is dm_crypt-0, the LUKS UUID is in /etc/crypttab):

sudo cp /etc/crypttab /etc/crypttab.bak
sudo sed -i 's/none luks$/none luks,tpm2-device=auto/' /etc/crypttab
cat /etc/crypttab     # → dm_crypt-0 UUID=… none luks,tpm2-device=auto
sudo update-initramfs -u
sudo reboot
  • Boots straight to login, no passphrase prompt = ✅ TPM auto-unlock works (unattended reboot achieved — VERIFIED on this unit 2026-06-23).
  • Still prompts = PCR mismatch; type the passphrase (NOT locked out), then retry with --tpm2-pcrs=0. The password slot + crypttab.bak make this fully reversible.

Re-seal runbook: a BIOS update / Secure Boot change alters PCR 7 → the TPM refuses → boot falls back to the passphrase prompt (not a brick). After such a change, re-run step 4's systemd-cryptenroll --wipe-slot=tpm2 --tpm2-device=auto --tpm2-pcrs=7 /dev/sda3 to re-bind.

5. GRUB password — EDIT-ONLY (VERIFIED 2026-06-23)

Closes the init=/bin/bash / systemd.unit=rescue.target local-root hole: without it, anyone at the keyboard presses e at the GRUB menu, edits the kernel cmdline, and boots to a root shell with no login. The PCR-7 TPM seal does NOT cover this — editing the GRUB cmdline doesn't change PCR 7 (Secure Boot policy), so the TPM still releases the key and the attacker lands on the decrypted disk. This is the specific countermeasure for the threat-model. Use edit-only mode (--unrestricted) so the box still boots UNATTENDED — the password is required only to EDIT entries, never to boot.

grub-mkpasswd-pbkdf2          # enter a password (twice) → copy the grub.pbkdf2.sha512.* hash

Add the superuser (paste YOUR hash) to the end of /etc/grub.d/40_custom:

set superusers="admin"
password_pbkdf2 admin grub.pbkdf2.sha512.10000.<YOUR_HASH>

Make menu entries bootable WITHOUT the password (edit-only) — in /etc/grub.d/10_linux, set the active CLASS= line to include --unrestricted:

CLASS="--class gnu-linux --class gnu --class os --unrestricted"

Regenerate + VERIFY BOTH HALVES landed in the real config BEFORE rebooting (a GRUB misconfig means a rescue-USB recovery):

sudo update-grub
sudo grep -c "password_pbkdf2" /boot/grub/grub.cfg   # want ≥1 (password present)
sudo grep -c "unrestricted"    /boot/grub/grub.cfg   # want ≥1 (entries bootable w/o password)
sudo reboot

✅ VERIFIED on this unit: boots straight to login (no GRUB prompt, TPM still auto-unlocks) AND pressing e at the menu prompts for admin + password. Store the GRUB password off-machine (alongside the LUKS passphrase).

OS hardening on the first unit is now COMPLETE: LUKS FDE + TPM auto-unlock (PCR 7) + Secure Boot (Deployed) + GRUB edit-lock.

5c. OS user model — admin vs operator (VERIFIED 2026-06-23)

The OS has TWO roles and they must be different identities (threat-model: the operator is the adversary). Create a dedicated admin (real password, sudo, NO auto-login) and keep the operator as an auto-login, UNPRIVILEGED account.

sudo adduser admin && sudo usermod -aG sudo admin
# VERIFY in a second session: log in as admin → `sudo whoami` prints root — BEFORE the next step:
sudo deluser <operator> sudo            # demote the auto-login operator
groups <operator>                        # confirm: no 'sudo'

⚠ Order matters: confirm the new admin's sudo works before demoting the operator, or you lock yourself out. Keep auto-login on the OPERATOR, not admin. Leave root password disabled (Ubuntu default) — admin+sudo IS the root path; enabling root adds risk, no gain.

Strip latent escalation groups from the operator: sudo deluser <operator> lxd (lxd group = launch a privileged container that mounts host / as root — undoes the no-sudo hardening) and lpadmin (printer admin, unneeded). And NEVER add the operator to docker (also root-equivalent).

5b. Further hardening (TODO — not yet done)

  • Key-based SSH only (disable password auth) if SSH is enabled at all. Routine ops no longer need SSH — Komodo Periphery (§7) drives deploys + gives a container terminal over the mesh — so SSH can be locked down hard or disabled, leaving the mesh + Komodo as the management path.
  • No/locked-down desktop + kiosk autostart — single-purpose; the operator never reaches a shell (desktop-shell-tauri).
  • Consider moving the host event-signing key into the TPM (non-extractable) — tpm, open-questions #12.
  • sudo apt autoremove the leftover old kernel once the new one is proven.

6. Runtime — Docker engine (VERIFIED 2026-06-23)

Install Docker Engine + compose (as admin). NB Ubuntu 26.04 codename is resolute, which download.docker.com may not yet publish — pin the repo line to noble, OR use Ubuntu's docker.io. Add only admin to the docker group (root-equivalent — NEVER the operator).

This gives the appliance the engine. How the stack gets ONTO it is step 7 — and as of 2026-06-27 the primary path is Komodo (remote, no-SSH), not a hand-copied dir. The manual docker compose flow survives as a break-glass fallback (§7c).

7. Deploy the stack — Komodo Periphery (PRIMARY, 2026-06-27)

The booth is driven by a central Komodo Core over the NetBird mesh. The appliance runs a small Periphery agent that dials out to Core; Core then deploys the same compose files. No inbound port on the booth, no SSH for routine ops. Full rationale + threat model: fleet-deployment-komodo. Verified end-to-end on the first booth (park-buzi) 2026-06-27.

7a. Install Periphery (on the booth, as admin)

Prereq: the booth is on the NetBird mesh and can reach Core's reverse-proxy URL (https://komodo.infra.msai.al).

  1. In Core: Settings → Onboarding → + New Onboarding Key (Name = the booth, e.g. park-buzi; Expiry ~1 day; Pre-Existing Key empty). Copy the one-time O-… key. Single-use — delete it after the agent connects.
  2. On the booth, install Periphery in user mode (runs as admin, who is in docker; NO root daemon; outbound → opens no inbound port):
curl -sSL https://raw.githubusercontent.com/moghtech/komodo/main/scripts/setup-periphery.py | python3 - --user \
  --core-address="https://komodo.infra.msai.al" \
  --connect-as="park-buzi" \
  --onboarding-key="O-…"
sudo loginctl enable-linger admin     # so the user service starts at boot without a login
  • --connect-as is the Server name in Core — unique, stable, site-meaningful (the fleet's primary key). Booth #2 = a different name (e.g. park-durres); never reuse one.
  • --core-address is Core's reverse-proxy URL (the URL you load the Core UI at over the mesh), NOT :9120 — Core's container port 9120 is exposed-not-published; the agent reaches it through the proxy. (Gotcha #7 below.)
  • Config lands at ~/.config/komodo/periphery.config.toml. The key field is core_address (singular); root_directory must be a path admin can write (user-mode default is fine — a /etc/komodo default from a system install would Permission denied for the user service).

Verify: systemctl --user status periphery → active; the server park-buzi appears and goes OK/green in Core → Servers. Then delete the onboarding key.

7b. Deploy the Stack (in Core — by hand once, then code)

Add Registry Account + Git Account for git.infra.msai.al (user komodo, tokens) in Core so Periphery can clone the repo AND pull the private images. Two distinct credential types — the git clone working does NOT imply the image pull is authed (gotcha #8). Per-booth secrets (park_<booth>_jwt_secret, park_<booth>_event_signing_key — distinct values, openssl rand -hex 32) live in Core's Variables/Secrets store, referenced from the Stack as [[…]].

Create a Stack (UI → Stacks → New), name = the booth (park-buzi):

  • Server: park-buzi · Source: repo mca/parking_solution, branch dev, files docker-compose.yml + docker-compose.prod.yml · Registry account: komodo (else the pull is anonymous → no basic auth credentials).
  • Environment (Komodo writes this to a .env on the booth at deploy, substituting [[…]]):
REGISTRY=git.infra.msai.al/mca/parking_solution
TAG=dev                                  # moving tag (staging). PIN to dev-<sha> for a live booth.
COOKIE_SECURE=0                          # CRITICAL on plain-http or the auth cookie never sends → no login
VISION_ENABLED=1
WS_ALLOWED_ORIGINS=                      # browser at the booth URL is same-origin; leave empty (the
                                         #   Tauri desktop app needs its origin here — separate task)
JWT_SECRET=[[park_buzi_jwt_secret]]
EVENT_SIGNING_KEY=[[park_buzi_event_signing_key]]

Deploy → Periphery pulls + compose ups. All containers (proxy/Caddy, server, vision) green. Seed the FIRST admin (DB starts empty → nobody can log in until this runs; idempotent) via Komodo's terminal on the server container (no SSH):

docker exec -it -e ADMIN_USER=admin -e ADMIN_PASS='<strong-pw>' \
  park-buzi-server-1 node scripts/seed-admin.mjs

Secrets-on-disk note. The generated .env lands on the booth with cleartext secrets (compose needs real values). That's why the disk is LUKS-encrypted (§3–4) and keys are per-booth — the encryption is the control, and a single-booth compromise leaks only that booth's key. See fleet-deployment-komodo (the EVENT_SIGNING_KEY-in-Core blast-radius caveat; ATECC608 is the intended long-term signer).

The repo's komodo/resources.toml mirrors the working Stack. Pointing a Core ResourceSync at it makes the fleet git-managed: booth #N is a copy-pasted [[stack]] block; an image bump is a one-line TAG= edit + push + Execute; every change is an auditable commit; a rebuilt Core re-creates everything from the file. Keep the sync Unmanaged + Delete-Unmatched OFF until trusted. An empty diff / disabled Execute = the file already matches the live Stack (success, not an error). See komodo/README.md and fleet-deployment-komodo.

7c. Break-glass — manual compose (mesh/Core down)

When the mesh or Core is unreachable, the same compose files run locally via scripts/booth.sh (or raw docker compose). Needs a local .env and a docker login git.infra.msai.al (a read-only package token). This is the FALLBACK, not the routine path:

docker login git.infra.msai.al
ENV=prod ./booth.sh config   # dry-run the merged env
ENV=prod ./booth.sh up

booth.sh runs from wherever it sits next to the compose files (the booth deploys them flat, e.g. /opt/parking_systems/). See container-deployment.

Healthy startup + web-access

Healthy logs: vision Initialized LicensePlateDetector … with NO "Downloading" (baked weights), server [migrate] done → SPA static serving enabled → Server listening. The transient vision-service -> offline at boot then -> ready (fast_alpr) ~8s later is normal (monitor polls before vision finishes loading). Reach the UI at http://<name-or-ip>/ (Caddy on :80).

Web-access gotchas (all fixed in the images/compose — see container-deployment "Web access"): the SPA uses a RELATIVE /api base (works from any host; do NOT bake a domain) + a Caddy proxy gives the clean port-80 URL; the domain (parksystems.msai.al) is pointed at the booth's LAN IP via hosts/DNS ON-SITE, never an image rebuild. The Tauri desktop app is hardcoded to localhost:3000 (CSP + endpoints) and can't reach a remote booth without code changes — a browser works; the desktop app is a separate workstream.

Quick-reference: the gotchas, in order they bit us

  1. Ventoy USB → 0x1A Security Violation under Secure Boot → flash the ISO directly instead.
  2. 7070 BIOS has no "View Key" → can't inspect db; the live-USB boot IS the verification.
  3. Installer "hardware-backed encryption" → PCR_UNUSABLE/dbt → use passphrase LUKS + manual seal.
  4. Bind TPM to PCR 7 only, not a multi-PCR set (kernel updates churn 4/8/9 → passphrase every boot).
  5. Always keep the password slot + an off-machine copy of the passphrase (TPM is never the only key).
  6. GRUB password MUST be edit-only (--unrestricted on entries) or it prompts on EVERY boot → breaks unattended reboot. Verify grep -c unrestricted /boot/grub/grub.cfg ≥1 before rebooting.

Komodo deploy gotchas (2026-06-27)

  1. Periphery core_address is Core's reverse-proxy URL (https://komodo.infra.msai.al), NOT 100.x:9120. Core's 9120 is exposed-not-published (docker ps shows 9120/tcp with no ->) → a direct dial gets Connection refused. Ping/SSH working over the mesh does NOT mean :9120 is reachable.
  2. Git auth ≠ registry auth. The repo cloning fine does not mean image pull is authed — they're separate Komodo credentials. A blank registry account on the Stack → anonymous pull → no basic auth credentials. Set the Stack's Registry Account (komodo).
  3. User-mode Periphery + /etc/komodo root_directory = Permission denied writing the agent key. User-mode (runs as admin, no root daemon) must keep root_directory under $HOME.
  4. The config key is core_address (singular). And --core-address derives wss:// from https:// — if Core were plain-HTTP you'd need http:// (→ ws://).
  5. ResourceSync Execute disabled + file shown clean in Info = empty diff = already in sync (success). Execute only enables when the file and Core diverge (e.g. you edit TAG).