Files
parking_solution/wiki/concepts/wsl-dev-networking.md
T
julian 28bd838696
Build desktop / desktop (push) Successful in 5m5s
CI / check (push) Successful in 47s
Build & push images / images (push) Successful in 2m59s
docs(wiki): merchant validations settled + as-built; scan input decided (camera paths postponed)
validation-discounts: driving cases → the settled validation-only model (all
money/paper at the booth) → setup UX/storage/RBAC → full as-built record.
DECIDED: merchant stations scan with a USB/HID barcode scanner on the
web/desktop app (hand-keying + Luhn as fallback); POSTPONED with analysis:
web getUserMedia scanning (secure-context TLS prerequisite on the LAN +
Code128-via-camera weakness → QR-on-ticket first) and a Tauri v2 Android
merchant app (native ML Kit scanning; Android build/sideload overhead +
configurable-server-URL prerequisite). Also: wsl-dev-networking gains the
mirrored-mode gotcha where a Windows-side listener makes a port EADDRINUSE
inside WSL while invisible to ss — Vite auto-increments and tauri dev's fixed
devUrl waits on the wrong port.

Claude-Session: https://claude.ai/code/session_01YYkpEsLmoQPaize5ec3oUm
2026-07-13 19:50:09 +02:00

130 lines
7.7 KiB
Markdown

---
type: reference
tags: [parking, dev-environment, networking, wsl, troubleshooting]
sources: []
updated: 2026-06-15
---
# WSL2 Dev Networking (for device testing)
> Dev-environment note, not product architecture. Recorded because reaching real
> hardware (the [[uhppote-controller]]) from a dev box running under **WSL2** took
> significant debugging. If you test devices from WSL, read this first.
## The problem
By default WSL2 uses **NAT networking**: the Linux VM sits on its own virtual subnet
(e.g. `172.x`), not the Windows host's LAN. Consequences for device work:
- **UDP broadcast (UHPPOTE discovery) cannot leave the VM** — a `get-devices` broadcast gets
`EACCES` / never reaches a controller on the physical LAN. The device is reachable from
*Windows* but not from *inside WSL*.
- Even unicast to a LAN device may not route, depending on setup.
## The fix: mirrored networking
Switch WSL to **mirrored** mode so it shares the Windows host's interfaces (and thus the real
LAN). Requires **Windows 11 22H2+** and **WSL ≥ 2.0**.
`%UserProfile%\.wslconfig` (create it; it doesn't exist by default):
```ini
[wsl2]
networkingMode=mirrored
firewall=false # Windows Firewall otherwise filters WSL traffic (can drop UDP replies)
[experimental]
hostAddressLoopback=true # host <-> WSL over the host's IP
```
Apply: in **PowerShell** `wsl --shutdown`, wait ~10 s, reopen WSL. Verify with `ip -4 addr` —
interfaces should now show the **real LAN subnet** (e.g. `10.0.10.x`) instead of `172.x`.
(Microsoft recommends editing via the **WSL Settings** GUI rather than the file by hand.)
> `wsl --shutdown` kills the dev servers — restart `pnpm dev` afterward.
## After mirrored mode: app-level gotchas that remained
Mirrored networking is necessary but **not sufficient** — these still bit us:
- **Multiple interfaces.** Mirrored WSL exposes *all* host NICs (LAN, Tailscale/CGNAT `100.x`,
docker bridges). UHPPOTE discovery must broadcast on **every** subnet, not the first one — see
[[device-discovery]].
- **Subnet-directed broadcast** (`10.0.10.255`, not `255.255.255.255`) — the lib won't enable
`SO_BROADCAST` otherwise. See [[device-discovery]].
- **`localhost` → IPv6 first.** `localhost` resolves to `::1`, but the backend binds IPv4
(`127.0.0.1`). Node's Vite proxy can stall on the v6 attempt before falling back — point the
proxy at `127.0.0.1` explicitly. (See [[local-dev-workflow]].)
- **Windows-side listeners collide with WSL binds — INVISIBLY (2026-07-13).** Under mirrored
mode, a process listening on the WINDOWS side makes the same port `EADDRINUSE` inside WSL,
but it never appears in Linux `ss`/`lsof` — the port looks free yet won't bind. Bit us as
"tauri dev: Could not connect to http://localhost:5173 after 180s": a DIFFERENT React app's
dev server running on the Windows side held `::1:5173`, so the WSL Vite silently
auto-incremented to 5174 while Tauri's `devUrl` is the FIXED string `http://localhost:5173`
in `tauri.conf.json` (it cannot follow the auto-increment). Diagnose from WSL with
`powershell.exe -NoProfile -Command "Get-NetTCPConnection -LocalPort 5173 -State Listen"`
(then `Get-Process -Id <OwningProcess>`); kill with `taskkill.exe /PID <pid> /F`. Guard:
`strictPort: true` in the web `vite.config` so the mismatch fails in a second with a clear
error instead of a 3-minute hang on the wrong port.
## Multi-subnet source-address trap (the "ARP works but ping/TCP dies" bug)
Field devices arrive **statically configured on assorted `/24`s** by whoever installed them last
(e.g. a camera on `10.0.10.121`, a printer on `10.0.10.6`, others on `192.168.1.x`). The host
copes by carrying **one IP per device subnet on a single NIC** (this is correct — you do **not**
need a NIC per subnet). But stacking subnets on one interface exposes a Linux source-selection
trap:
- Connected routes come up as `proto kernel scope link` **with no preferred source**. With two
such subnets on one NIC, the kernel may pick the **wrong source address** — e.g. sourcing
traffic to `10.0.10.121` from `192.168.1.123`.
- Symptom is baffling: **ARP resolves and the neighbor shows `REACHABLE`** (L2 is fine, source
address is irrelevant to ARP) while **every ping and TCP connect times out** (replies have a
wrong/unroutable source → dropped, possibly by uRPF). Looks like "the device is down / the whole
subnet is unreachable" when nothing is actually broken.
- **Diagnose:** `ip route get <device-ip>` shows the chosen `src` — if it's an address on a
*different* subnet, that's the bug. Confirm by forcing the right source:
`ping -I <correct-src> <device-ip>` (or `curl --interface <correct-src> …`) — instant replies.
- **Fix (runtime):** pin the preferred source on the connected route, per subnet:
`sudo ip route replace <subnet>/24 dev <nic> proto kernel scope link src <correct-host-ip> metric <m>`
(use `replace`, not `change` — `change` errors `RTNETLINK: No such file` if the route isn't up
yet). Do **not** delete the other subnet's address unless it's genuinely unwanted — you need all
of them to reach all the devices.
- **Fix (permanent, this box):** `deploy/wsl-fix-route-source.sh` + `deploy/parking-net.service`.
The script walks each `proto kernel scope link` route on the NIC and pins `src` to THIS host's own
address in that same subnet — **no hardcoded IPs**, so it also covers future device subnets; it's
idempotent, preserves the route metric, and tolerates a missing route. The systemd unit (oneshot,
`enabled`) reapplies it on every WSL boot — which is the point, since `wsl --shutdown` otherwise
wipes the runtime fix (mirrored mode re-clones the Windows addresses fresh each boot, see below).
Install once: copy the unit to `/etc/systemd/system/`, `systemctl enable --now parking-net`.
Gotchas hit while building it: `network.target` is too early for mirrored-mode addresses (the
script waits up to 15s for a route to appear); and it must NOT `set -e` or one failed `ip` call
aborts the whole boot fixer.
> **Root cause is on the Windows side.** Mirrored mode clones the Windows host NIC's addresses into
> Linux at every boot, so the stray `192.168.1.x` lives on Windows — the truly permanent fix is to
> remove/reconfigure it there (or set `SkipAsSource`/interface metric). The systemd hook is the
> self-contained Linux-side answer that needs no Windows changes.
Verified on hardware (2026-06-15): after the hook, `10.0.10.121` pings and the real [[lpr-camera]]
Hikvision driver pulls a snapshot with **no** source-forcing (`localAddress` becomes optional).
## On the real appliance: multi-subnet is a deployment config, not a WSL hack
Production is a **dedicated hardened Linux appliance** ([[disk-os-hardening]]), so the WSL story
above is dev-only. The device-subnet problem persists, though, and is solved the same way at the
OS level: the appliance NIC carries **one address per device subnet**, each connected route with a
pinned `src`, made persistent (systemd-networkd / netplan). Per the threat model this still rides
on **[[network-isolation]]** — device subnets are isolated segments reachable only by the host.
The long-term clean answer is to **re-IP the devices onto one planned parking-system subnet** at
install so the host needs only one address; the multi-subnet config is what you run until then.
## Alternative if you can't use mirrored mode
Windows 10 / old WSL can't do mirrored mode. Options: run the **backend natively on Windows**
(shares the LAN), or use **unicast by IP** instead of broadcast discovery (target the controller's
known IP — the driver supports an explicit host). On the real **appliance** (a dedicated hardened
Linux box, [[disk-os-hardening]]) none of this applies — it's bare-metal on the device VLAN
([[network-isolation]]).