f7a262ac9a
Build & push images / images (push) Successful in 6m31s
apps/trainer (parking-trainer): inspect / train / evaluate / publish. Reads the wash collector's SQLite + crops read-only off its volume; time split (validation = newest slice); thin classes dropped; damped class weights; `features` mode (frozen ImageNet backbone, on-disk feature cache, seconds to retrain) and `finetune` mode (light augmentation). CPU-only torch from PyTorch's wheel index. ONNX export checked against the torch model; NO model file below the validation floor (exit 3, report still written); exit 2 = not enough labels. `evaluate` scores a shipped model on labels reviewed after training + the unlabelled pile; `publish` PUTs a version folder to a Gitea generic package. Light core deps; the `train` extra is heavy — CI syncs without it, torch tests skip. apps/vision: BodyTypeClassifier (bodytype.onnx + sidecar = the preprocessing contract: crop margin, input size, RGB 0-255, normalisation inside the graph) and RefinedVehicleDetector over YOLOX — refines only `car` or a class the classifier trained on, min-confidence, `detector_class` on the result; path set but no file = phase B off without an error; a broken file is a health detail. models/bodytype.version (tracked, empty) pins the published version the Dockerfile fetches at build (BuildKit secret; a pin that cannot be fetched fails the build). Verified: a trainer model gives identical probabilities inside the vision service; both images built and smoke-tested. Delivery: parking-trainer image in build-images.yml, the `trainer` compose profile on the collector stack (CPU, read-only data, TRAINER_OUT), commented TRAINER_OUT/PUBLISH_TOKEN in the wash-collector stack, .dockerignore for both Python contexts, trainer deps synced in CI. Wiki: bodytype-classifier-training rewritten as built (+ one fleet model not per site, secrets/access, where the crops live), opencv-anpr-service §Phase B, vision-review-outbox, vision-service-packaging, fleet-deployment-komodo, index, log. Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
51 lines
2.8 KiB
YAML
51 lines
2.8 KiB
YAML
# The Car Wash REVIEW COLLECTOR — deployed on the REVIEWER's host (art-docker-station), NOT
|
|
# on a booth. Its own Komodo stack ("wash-collector" in komodo/resources.toml) points at this
|
|
# file alone, so nothing here reaches a booth and nothing of the booth stack reaches this
|
|
# host. See wiki/concepts/vision-review-outbox.md.
|
|
#
|
|
# Reachability: booths POST to /ingest over the Netbird overlay only. Bind the published
|
|
# port to the host's OVERLAY address (COLLECTOR_BIND), never 0.0.0.0 on a host that also
|
|
# has a public interface. The Netbird policy should allow booths → this host:8090 and
|
|
# nothing else on it.
|
|
|
|
services:
|
|
collector:
|
|
image: ${REGISTRY:-git.infra.msai.al/mca/parking_solution}/parking-collector:${TAG:-dev}
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${COLLECTOR_BIND:-127.0.0.1}:8090:8090"
|
|
environment:
|
|
# "<boothId>:<token>" pairs — one per booth, the booth's CARWASH_REVIEW_TOKEN under its
|
|
# pseudonymous CARWASH_REVIEW_BOOTH_ID. A Komodo secret reference in the stack env.
|
|
COLLECTOR_BOOTH_TOKENS: ${COLLECTOR_BOOTH_TOKENS:?set COLLECTOR_BOOTH_TOKENS in the stack env}
|
|
# The single reviewer login (HTTP Basic over the overlay).
|
|
COLLECTOR_REVIEWER_USER: ${COLLECTOR_REVIEWER_USER:-reviewer}
|
|
COLLECTOR_REVIEWER_PASS: ${COLLECTOR_REVIEWER_PASS:?set COLLECTOR_REVIEWER_PASS in the stack env}
|
|
LOG_LEVEL: ${LOG_LEVEL:-info}
|
|
volumes:
|
|
- collector-data:/data
|
|
|
|
# Phase B trainer — a ONE-OFF JOB on this host's CPU, not a service (profile "train": it
|
|
# only runs when asked). Reads the collector's SQLite + crops straight off the same volume
|
|
# (read-only), writes a versioned model folder under TRAINER_OUT on the host. CPU-only
|
|
# PyTorch: the Xeon E3-1225 v5 trains a few thousand crops in minutes (features mode) to an
|
|
# hour (full fine-tune) — see wiki/decisions/bodytype-classifier-training.md. If a modern GPU
|
|
# ever lands in the host, add an nvidia device reservation here; the trainer picks up CUDA.
|
|
#
|
|
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer inspect
|
|
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer train --min-accuracy 0.85
|
|
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer evaluate --model /out/<version>/bodytype.onnx
|
|
# docker compose -f docker-compose.collector.yml --profile train run --rm trainer publish /out/<version> --url <gitea generic package url>
|
|
trainer:
|
|
image: ${REGISTRY:-git.infra.msai.al/mca/parking_solution}/parking-trainer:${TAG:-dev}
|
|
profiles: ["train"]
|
|
environment:
|
|
# Only `publish` needs it: a Gitea token with package:write for the model's generic package.
|
|
TRAINER_PUBLISH_TOKEN: ${TRAINER_PUBLISH_TOKEN:-}
|
|
volumes:
|
|
- collector-data:/data:ro
|
|
- ${TRAINER_OUT:-./models}:/out
|
|
|
|
volumes:
|
|
collector-data:
|