Files
parking_solution/wiki/decisions/fleet-deployment-komodo.md
T
julian 9918f278b2
CI / check (push) Successful in 36s
feat(deploy): Komodo fleet deployment — resources.toml + decision
Adopt Komodo Periphery (over the NetBird mesh) as the booth fleet control plane,
superseding SSH-and-booth.sh. The booth runs the SAME compose files; Komodo Core
drives them remotely. booth.sh is demoted to a break-glass local fallback.

- komodo/resources.toml mirrors the working park-buzi Stack (built by hand in the
  Core UI, then exported to TOML — field names match the running v2.2). Stack-only:
  servers are created by the agent onboarding OUTBOUND (one-time onboarding key →
  Periphery self-registers, auto-rotating keys, booth opens no inbound port), so
  there is no [[server]] block. Per-booth secrets via [[...]] refs to Core's store.
- komodo/README.md + .env.komodo.example document the flow and the hard rules
  (no webhook; onboarding/outbound/mesh-only; per-booth unique secrets; never
  down -v the ledger volume).
- wiki/decisions/fleet-deployment-komodo.md records the decision + threat-model
  analysis (Periphery is a root agent → mesh-bound; EVENT_SIGNING_KEY-in-Core is a
  fraud-root blast radius until ATECC608 signs; Core is now Tier-0; GPL-3.0 is fine
  as external ops tooling). container-deployment reframed (booth.sh = fallback);
  index + log updated.

Verified end-to-end against a real booth (park-buzi): onboarded OK, Stack deployed,
all containers green, admin seeded.

Claude-Session: https://claude.ai/code/session_01Xcm6ikLgGoCxxHrxtjkk5V
2026-06-27 12:14:04 +02:00

8.4 KiB

type, tags, sources, updated, status
type tags sources updated status
decision
parking
deployment
fleet
komodo
netbird
offline-first
threat-model
2026-06-27 settled

Fleet deployment — Komodo Periphery over a NetBird mesh

How the parking appliance is deployed and managed at fleet scale, superseding the single-box, SSH-and-booth.sh model. The image build/tag/registry pipeline (container-deployment) is unchanged — this decides only the control plane that drives those same compose files onto many booths. Settled 2026-06-27.

The problem booth.sh couldn't solve

[[container-deployment|scripts/booth.sh]] is a thin wrapper over docker compose -f base -f prod --env-file .env. It works for one appliance you can get a shell on, but as the fleet grows (the stated direction is many/growing sites) it gives us none of:

  • Remote, no-SSH operation — an update means someone gets a root shell on the booth.
  • A fleet view — which booth runs which dev-<sha>, which is healthy/offline.
  • A deploy audit trail — who deployed what, when.
  • One-click rollback to a previous immutable dev-<sha>.

These are exactly the gaps a deployment controller fills. We already run every prerequisite (a Komodo Core, a NetBird zero-trust mesh, the Gitea registry), so the marginal cost is low.

Decision

Adopt Komodo Periphery on each appliance, driven by the existing Komodo Core over the NetBird mesh. Keep the compose files and the container-deployment verbatim — Komodo consumes them as a Stack; it does not replace them. booth.sh is demoted to a break-glass local fallback for when the mesh/Core is unreachable.

Gitea push ─▶ build-images.yml ─▶ registry (parking-server:dev-<sha>, parking-vision:dev-<sha>)
                                          │
Komodo Core (off-site) ──── NetBird mesh ─┼─▶ Periphery @ booth-A  ─▶ docker compose up (pinned sha)
   • fleet table / history / rollback     ├─▶ Periphery @ booth-B
   • per-booth secret injection           └─▶ Periphery @ booth-C …
   • NO deploy webhook (manual + pinned)

The three load-bearing choices (settled with the user 2026-06-27)

  1. Fleet size: many/growing. Komodo is treated as load-bearing infrastructure, not a convenience. This is what tips the decision away from "SSH-over-NetBird + a playbook".
  2. Deploy trigger: always manual + pinned. No deploy webhook on a booth Stack. A human deploys a specific immutable TAG=dev-<sha> from Core. This preserves the determinism we chose when pinning the booth tag (a moving :dev auto-redeploying a production booth is the surprise we explicitly rejected). A staging booth MAY track :dev; a production booth never does.
  3. Secrets: Komodo-managed (per-booth, unique). Core's secret store injects JWT_SECRET and EVENT_SIGNING_KEY into the Stack at deploy. This scales (no SSH-to-N-booths to rotate a key) — but see the threat-model tension below; the keys MUST be distinct per booth.

Why this is safe (against the project's two forces)

Offline-first (offline-first) — Core is orchestration, never a runtime dependency

The booth must run fully when the mesh is down. Komodo's agent model satisfies this: Periphery + the local containers keep operating if Core is unreachable; we lose remote management until the mesh returns, not operation. There must be no runtime path from booth operation to Core — Core only deploys. (Periphery's own liveness is irrelevant to entry/ exit; the Fastify server and SQLite ledger run independently of it.)

Threat model — the adversary is the booth operator (threat-model)

This is the sharp edge, and the reason this page is explicit rather than a footnote.

  • Periphery is a root-capable remote-exec agent on the appliance. If the operator compromises the box, the agent is a lever. Mitigations: bind Periphery only to the NetBird interface (never 0.0.0.0), enforce its passkey + TLS, and fold the agent into the disk-os-hardening surface. It is part of the trusted computing base now.
  • EVENT_SIGNING_KEY is the anti-fraud root. It signs the [[append-only-event-chain| append-only ledger]] — the control between us and a booth operator forging entry/exit events. Holding it in Core means a Core compromise can forge any booth's ledger that shares a key. Two mitigations make central management acceptable:
    • Per-booth, unique keys. Never reuse a signing key across sites, so a single leak taints one booth, not the fleet.
    • The atecc608 is the real long-term signer. The EVENT_SIGNING_KEY HMAC is the interim mechanism; once the secure element signs the chain, the key in Core stops being the fraud root. Tracked in open-questions.
  • Core becomes a Tier-0 asset. It now holds login + ledger keys for the whole fleet, so it must be hardened to the booths' bar: Komodo API bound to the NetBird mesh only, never a public interface; access-controlled; backed up.

Licensing — Komodo is GPL-3.0, and that's fine here

The hard MIT/Apache/BSD constraint (technology-stack) is about shipped app dependencies (code we distribute/link). Komodo is external ops tooling we self-host and don't distribute, so its GPL-3.0 does not taint the product — exactly like the [[vision-service|AGPL ANPR exception]] reasoning (a separate process / external boundary, not a linked dependency). Noted here so it isn't re-litigated.

What lives where

Concern Where Notes
Image build + tags Gitea CI (container-deployment) unchanged: :dev moving + :dev-<sha> immutable
Compose files the repo + on the booth unchanged base + docker-compose.prod.yml
Stack / deploy definition Komodo Core git-synced from komodo/ (infra-as-code)
Which sha is deployed Komodo Core, manual TAG=dev-<sha>, pinned, no webhook
JWT_SECRET, EVENT_SIGNING_KEY Komodo Core secret store per-booth, unique
COOKIE_SECURE=0, TAG, REGISTRY Komodo Stack env per-environment
Registry pull creds Komodo Core so Periphery can pull from Gitea
Local break-glass booth.sh + a local .env mesh-down fallback only

Setup outline

On each appliance (after appliance-provisioning):

  1. Install Komodo Periphery (binary or container), bound only to the NetBird interface; set its passkey/TLS.
  2. Point its compose/stack dir at /opt/parking_systems/ (the existing files).
  3. Keep booth.sh + a minimal local .env (no real secrets) as break-glass.

In Komodo Core:

  1. Add the booth as a Server, address = its NetBird IP (mesh, not LAN/WAN).
  2. Define the Stack = base + docker-compose.prod.yml, env from Core's secret store, secrets per booth.
  3. No deploy webhook on the booth Stack — deploys are manual; set TAG=dev-<sha> explicitly.
  4. Add Gitea registry creds so Periphery can pull.
  5. Sync the Stack/Server definitions from the repo's komodo/ directory (infra-as-code: komodo/resources.toml + README) so the control plane is itself reviewable + version-controlled.

Open / not yet done

  • Per-booth secret generation + rotation flow — how a new site's unique EVENT_SIGNING_KEY is generated and registered in Core (vs. on-site openssl rand). Tie-in: open-questions JWT-key item.
  • ATECC608 as the signer supersedes EVENT_SIGNING_KEY-in-Core as the fraud root — until then central secrets carry the blast-radius noted above.
  • Periphery hardening checklist folded into disk-os-hardening (interface binding, passkey, TLS, agent as TCB).
  • Staging vs production booth split (a staging booth on :dev with a webhook; production manual+pinned) — not yet modelled in komodo/.
  • Core backup / DR — Core is now Tier-0; its loss = no fleet management (operation unaffected, per offline-first). Backup story TBD.

Supersedes / relates

  • Supersedes the "SSH + booth.sh is the deploy mechanism" assumption in container-deployment (that page's build/tag/registry content stands; its booth.sh-as- primary-deploy framing is now the fallback). Cross-linked there.
  • Companion: the komodo/ infra-as-code sketch (in the repo, not the wiki), appliance-provisioning (what runs before Periphery), disk-os-hardening (the appliance's hardening surface).