feat(trainer): phase-B body-type classifier — trainer job on the collector host + the classifier stage on the booth
Build & push images / images (push) Successful in 6m31s

apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash
collector's SQLite + crops read-only off its volume; time split (validation = newest
slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet
backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light
augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against
the torch model; NO model file below the validation floor (exit 3, report still written);
exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after
training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package.
Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip.

apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract:
crop margin, input size, RGB 0-255, normalisation inside the graph) and
RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained
on, min-confidence, `detector_class` on the result; path set but no file = phase B off
without an error; a broken file is a health detail. models/bodytype.version (tracked,
empty) pins the published version the Dockerfile fetches at build (BuildKit secret;
a pin that cannot be fetched fails the build). Verified: a trainer model gives identical
probabilities inside the vision service; both images built and smoke-tested.

Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the
collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in
the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI.

Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site,
secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox,
vision-service-packaging, fleet-deployment-komodo, index, log.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
This commit is contained in:
2026-09-07 11:14:50 +02:00
parent f9cb973fe9
commit f7a262ac9a
41 changed files with 3797 additions and 76 deletions
+20 -17
View File
@@ -25,23 +25,26 @@ services:
volumes:
- collector-data:/data
# Phase B trainer — a one-off job on this host's GPU, NOT a service (profile "train": it
# only runs when asked: `docker compose --profile train run --rm trainer`). Reads the
# collector's export + crops straight off the same volume; writes the ONNX classifier the
# vision image then bakes in. The image/script are the next increment; this is the seam.
# trainer:
# image: ${REGISTRY:-git.infra.msai.al/mca/parking_solution}/parking-trainer:${TAG:-dev}
# profiles: ["train"]
# deploy:
# resources:
# reservations:
# devices:
# - driver: nvidia
# count: all
# capabilities: [gpu]
# volumes:
# - collector-data:/data:ro
# - ./models:/out
# Phase B trainer — a ONE-OFF JOB on this host's CPU, not a service (profile "train": it
# only runs when asked). Reads the collector's SQLite + crops straight off the same volume
# (read-only), writes a versioned model folder under TRAINER_OUT on the host. CPU-only
# PyTorch: the Xeon E3-1225 v5 trains a few thousand crops in minutes (features mode) to an
# hour (full fine-tune) — see wiki/decisions/bodytype-classifier-training.md. If a modern GPU
# ever lands in the host, add an nvidia device reservation here; the trainer picks up CUDA.
#
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer inspect
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer train --min-accuracy 0.85
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer evaluate --model /out/<version>/bodytype.onnx
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer publish /out/<version> --url <gitea generic package url>
trainer:
image: ${REGISTRY:-git.infra.msai.al/mca/parking_solution}/parking-trainer:${TAG:-dev}
profiles: ["train"]
environment:
# Only `publish` needs it: a Gitea token with package:write for the model's generic package.
TRAINER_PUBLISH_TOKEN: ${TRAINER_PUBLISH_TOKEN:-}
volumes:
- collector-data:/data:ro
- ${TRAINER_OUT:-./models}:/out
volumes:
collector-data: