Files
parking_solution/apps/vision
julian 20a3cb3e80 feat(vision): vehicle stage, phase A — YOLOX-S (Apache-2.0 ONNX) beside the plate recognizer
Fills /analyze vehicle.body_type + confidence (car / motorcycle / bus / truck from COCO,
mapped to the shared vocabulary) for the Car Wash desk's category suggestion
(venue-modules.md §Vehicle category from vision). Advisory: the operator decides, a
confident downgrade is flagged, nothing is gated on it.

- vision_service/vehicle.py: pure numpy/cv2 letterbox (pad 114, raw BGR), stride-grid
  decode, class-agnostic NMS, one vehicle per frame (the box holding the plate's centre,
  else the largest); YoloxVehicleDetector on onnxruntime CPU, 2 intra-op threads.
- recognizer.py: WithVehicle composes the stage over any plate recognizer (stub included);
  a failing stage yields vehicle=null + a "vehicle: …" note in /health.detail — never
  costs the plate read. model_version reads "<plate>+yolox:yolox_s.onnx@640".
- settings: VISION_VEHICLE_MODEL_PATH (unset = off), _INPUT_SIZE (640), _MIN_CONFIDENCE
  (0.4, the detector's floor; the flag threshold is site config).
- Dockerfile bakes yolox_s.onnx (best-effort curl at build; no network → stage off) and
  sets the path; compose forwards it (empty = off); .env.example documents it.
- Measured on four real dev entry frames (DS-2CD1047G3H, 2560×1440): car at 0.83–0.88 in
  ~240–330 ms; empty lane with a person → none.
- tests/test_vehicle.py: decode/NMS/pick/letterbox on synthetic tensors, the composition,
  and a missing-model /health. Wiki: opencv-anpr-service, venue-modules, log.

Claude-Session: https://claude.ai/code/session_01FWncR69HgGPuei1dLrW3cU
2026-09-06 19:53:11 +02:00
..

@parking/vision — host-side ANPR / vehicle-verification service

A separate process (Python + FastAPI) the Node backend calls over localhost HTTP with a camera snapshot, returning a licence-plate read (and, later, vehicle-attribute verification — the anti-plate-spoofing witness). Recognition is advisory, never the sole authority to open a barrier: if this service is down or unsure, the host falls back to the ticket path.

Lives inside the Turborepo at apps/vision/ but is not a JS package — Python deps are managed by uv/pyproject.toml; the package.json is a thin shim so turbo run lint/test includes it. See wiki/decisions/vision-service-packaging.md and wiki/entities/opencv-anpr-service.md.

Run

# from apps/vision/ — install the light core (boots in stub mode, no model downloads)
uv sync

# dev server with reload (or: pnpm --filter @parking/vision dev)
uv run uvicorn vision_service.app:app --reload --port 8089

# checks
uv run ruff check .
uv run pytest -q

Enable the real recognizer (fast-alpr)

uv sync --extra alpr                 # installs fast-alpr + onnxruntime (downloads model weights)
VISION_RECOGNIZER=fast_alpr uv run uvicorn vision_service.app:app --port 8089

Model weights (~11 MB: a YOLOv9 detector + CCT OCR) download on first use and cache under ~/.cache/open-image-models + ~/.cache/fast-plate-ocr — offline after that.

Quick test against an image (CLI, no HTTP)

uv run python -m vision_service.cli path/to/car.jpg          # or: pnpm --filter @parking/vision recognize -- car.jpg
uv run python -m vision_service.cli car.jpg --ocr cct-s-v2-global-model   # try another OCR model

Prints the parsed plate(s) + confidence + region as JSON. Confidence is the min of fast-alpr's per-character confidences (a plate is only as trustworthy as its weakest character). Example output on the fast-alpr test image: 5AU5341 (1.000) region "Czech Republic" in ~40 ms on CPU.

fast-alpr is MIT (YOLOv9 detector + CCT OCR on ONNX Runtime). Swap VISION_OCR_MODEL to the 40+ country European model to benchmark Albanian plates. For GPU/NPU, install onnxruntime-gpu / -openvino / -directml instead of onnxruntime.

API

  • GET /health → { status, recognizer, ready, model_version, detail? }
  • POST /analyze (body = raw image bytes, Content-Type: application/octet-stream) → { plate: {text, confidence, bbox}|null, plates[], vehicle: null, low_confidence, model_version, took_ms }

The Node side POSTs Snapshot.bytes directly (no multipart). vehicle is scaffolded but not yet populated — fast-alpr is plate-only; the vehicle stage (Job 2) is built later on the same runtime.

Config (env, prefix VISION_) — see .env.example

This service's env only. The Node server has its own VISION_* (apps/server/.env: VISION_ENABLED, VISION_URL, VISION_POLL_MS, …) — same prefix, separate process, separate .env. Don't merge them.

Var Default Meaning
VISION_RECOGNIZER stub stub (no models) or fast_alpr (real)
VISION_HOST 0.0.0.0 bind address — prefer 127.0.0.1 on the appliance (Node is the only caller)
VISION_PORT 8089 listen port (must match the server's VISION_URL)
VISION_DETECTOR_MODEL yolo-v9-t-384-license-plate-end2end fast-alpr detector
VISION_OCR_MODEL cct-xs-v2-global-model fast-alpr OCR (won the AL benchmark)
VISION_MIN_CONFIDENCE 0.5 below this → low_confidence=true

To use it from the booth: set VISION_ENABLED=1 on the server, run this service, then tick ANPR on a camera in the SetupWizard (the camera must also be bound to a barrier). The booth footer shows a Vision chip when enabled. Full config guide: wiki/entities/opencv-anpr-service.md.