Files
parking_solution/komodo
julian f7a262ac9a
Build & push images / images (push) Successful in 6m31s
feat(trainer): phase-B body-type classifier — trainer job on the collector host + the classifier stage on the booth
apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash
collector's SQLite + crops read-only off its volume; time split (validation = newest
slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet
backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light
augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against
the torch model; NO model file below the validation floor (exit 3, report still written);
exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after
training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package.
Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip.

apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract:
crop margin, input size, RGB 0-255, normalisation inside the graph) and
RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained
on, min-confidence, `detector_class` on the result; path set but no file = phase B off
without an error; a broken file is a health detail. models/bodytype.version (tracked,
empty) pins the published version the Dockerfile fetches at build (BuildKit secret;
a pin that cannot be fetched fails the build). Verified: a trainer model gives identical
probabilities inside the vision service; both images built and smoke-tested.

Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the
collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in
the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI.

Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site,
secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox,
vision-service-packaging, fleet-deployment-komodo, index, log.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
2026-09-07 11:14:50 +02:00
..

komodo/ — fleet deployment as code

Infra-as-code for the Komodo Core control plane that deploys the parking appliance to the booth fleet over the NetBird mesh. See wiki/decisions/fleet-deployment-komodo.md for the rationale, threat-model analysis, and the three settled choices (many/growing fleet · deploys are manual + pinned · secrets are Komodo-managed, per-booth).

This directory does not change how images are built or how the app runs — it's only the control plane. The booth still runs the same docker-compose.yml + docker-compose.prod.yml (container-deployment); Komodo just drives them remotely instead of someone SSH-ing in to run booth.sh.

Files

  • resources.toml — the Komodo resource definitions (Servers, Stacks, optional Builders/ Procedures), synced into Core via a ResourceSync. This is the reviewable, version- controlled source of truth for which booth runs what.
  • .env.komodo.example — the variables a Stack expects, documenting what comes from Core's secret store (per-booth JWT_SECRET / EVENT_SIGNING_KEY / BACKUP_KEY) vs. plain Stack env.

Promotion: dev → stage → main

Three tiers (see wiki/decisions/fleet-deployment-komodo.md):

  • dev — the working branch. CI builds :dev / :dev-<sha>. No booth deploys off it.
  • stage — what the staging booth (park-buzi) runs, to test the app in real-world conditions. When dev is confident-ready, merge dev → stage; CI builds :stage / :stage-<sha>; then bump TAG=stage-<sha> in resources.toml to that build and deploy from Core (manual, pinned — no webhook even on staging).
  • main — vetted production booths: :main-<sha>, manual + pinned. main only gets what survived staging.

The TAG=stage-<sha> in resources.toml is a pinned pointer: the moving :stage tag exists but we deploy the immutable sha so a booth runs a known image. Re-pin on each promotion.

How Core consumes this (one-time)

In Komodo Core, create a ResourceSync pointing at this repo + path (komodo/resources.toml), on the branch you manage from (e.g. main). Core reads the file and reconciles Servers/Stacks to match. Thereafter, a PR to this directory + a sync is how you change the fleet — no clicking.

Komodo's TOML schema evolves across releases. Treat resources.toml as a starting sketch: resources.toml mirrors the working park-buzi Stack (built by hand in the Core UI, then exported to TOML — so field names match the running Komodo version, v2.2). Import it into the sync Unmanaged first and review the diff; it should be ~empty against the live Stack.

How servers get created — NOT here

There is no [[server]] block in resources.toml. Servers are created by the Periphery agent onboarding outbound: in Core, create a one-time Onboarding Key (Settings → Onboarding), then install Periphery on the booth passing --onboarding-key + --core-address (Core's reverse-proxy URL, reached over the NetBird mesh) + --connect-as=<booth-name>. The agent self-registers, generates its own auto-rotating key pair (private key never leaves the booth), and connects outbound — the booth opens no inbound port. The sync owns only the Stack, which references the server by the name it onboarded as (server = "park-buzi"). See wiki/decisions/fleet-deployment-komodo.md.

Adding booth N

Copy the [[stack]] block, change name, server (its onboarded name), and the per-booth secret references ([[park_<site>_jwt_secret]], [[park_<site>_event_signing_key]], [[park_<site>_backup_key]]). Create those secrets in Core's store first. Set branch + TAG for the tier the booth runs (staging → stage / stage-<sha>; production → main / main-<sha>).

Hard rules encoded here (do not relax without updating the decision page)

  1. No deploy webhook on a booth Stack. Deploys are a human action; pin an immutable TAG=<branch>-<sha> (staging → stage-<sha>, production → main-<sha>). A moving tag on a booth is the non-determinism we rejected — and we hold that line even on the staging booth.
  2. Onboarding, outbound, mesh-only. Servers self-register via an onboarding key; Periphery connects outbound to Core's mesh URL and exposes no inbound port. Never a LAN/WAN address.
  3. Secrets are per-booth and unique. EVENT_SIGNING_KEY signs the anti-fraud ledger — one leak must taint one booth, never the fleet. Reference Core secrets by name; never inline a real value in this file (it's in git).
  4. Volumes preserved. The Stack must never run compose down -v — that would wipe the parking-data volume (the signed ledger). Komodo's "destroy" is gated for the same reason.