# Experiment 03: converged RadioMapSeer runs on the Mac GPU (MPS)

_Started 2026-09-29. Task: `docs/TASK-claude-code-01-ml-on-mps.md`. Code: `rcm_ml/`. Raw results: `runs/mps/**/test.json` (each has the commit hash, seed, config and wall time)._

## 1. Setup

- Apple M2 Max (12 CPU cores, 38-core GPU), 64 GB, macOS 26. Python 3.12 venv (`uv`), PyTorch 2.14, backend **MPS**. Native, not Docker. `train.py` exits if `torch.backends.mps.is_available()` is false. No `PYTORCH_ENABLE_MPS_FALLBACK`: nothing needed it.
- Data: RadioMapSeer zip (3.26 GB) → `data/rms/`. Cache: `python -m rcm_ml.prepare --out data/rms_cache --tx-train 80 --tx-eval 80 --workers 10`, **3 min**, 14 GB uint8 memmaps (train 40,080 / val 8,000 / test 7,920 samples, all 80 Tx per map).
- Verified from the data: antenna PNG row = 255 − json y; exactly one Tx pixel; buildings PNG ∈ {0, 255}; gains are uint8 0–255.

## 2. Two assumptions that were wrong (fixed before any training)

Both come from reading the official code, `github.com/RonLevie/RadioUNet` (`lib/loaders.py`, `RadioWNet_c_DPM_Thr2.ipynb`).

**2.1 The official split exists, and it is not ours.** The loader shuffles `arange(700)` with `np.random.seed(42)` and adds 1, so the dataset uses map ids **1–700** and leaves out **map 0** (our docstring said map 700). Index ranges are inclusive: train 0–500, val 501–600, test 601–699, i.e. **501 / 100 / 99 maps**. Added as `split(official=True)` (test checks it against the loader's code). The cache and every run in this document use it.

**2.2 The paper's RMSE is on floor-clipped targets.** `RadioUNet_c(thresh=0.2)` is the default: targets below gray 0.2 (−127.2 dB, the SNR = 0 threshold) are set to 0.2, then rescaled `(f − 0.2) / 0.8`. So one paper gray unit is 0.8 × 99.16 = **79.3 dB**. That is where the paper's "dB = 80 × gray" comes from, and it resolves the 80-vs-99.16 inconsistency noted in `01-literature-and-datasets.md`. **Consequence: RadioUNet_C's 0.0200 gray is ≈ 1.59 dB on the clipped scale, not 2.0 dB.** The goal-1 bar "≤ 2.3 dB" is kept, but the reference number is 1.6 dB.

What we now do:

- Train on clipped targets (`--thresh 0.2`, the paper protocol) for all variants, so that `geom`, `feats` and `hybrid` compete on the same objective. The B1 channel is put through the same clip and rescale.
- Report, on the test set:
  - `rmse_gray_paper`: comparable to the paper's table
  - `rmse_db_thr`: the same in dB (**primary metric**)
  - `rmse_db_thr_outdoor`: only pixels outside buildings
  - `rmse_db_raw`: against the raw PNG (the metric of experiment 02, which punishes not predicting sub-floor values)
  - `rmse_db_above_-127`: pixels whose true value is above −127 dB

**Why it matters for experiment 02:** its 3.3 / 4.3 dB numbers are "raw" RMSE at 128 px. They are not on the same scale as the paper's 2.0 dB (their clipped equivalent would be lower). The 02 conclusions are relative comparisons, so they stand, but "we are 1.3 dB from the paper" was measured against the wrong reference.

## 3. Throughput (Phase A)

`python -m rcm_ml.bench` → `runs/mps/bench.json`. Training step on synthetic tensors, batch 15, samples/s:

| size | U-Net base (params) | fp32 | fp16 autocast | fp32 + channels_last | fp16 + channels_last |
|---|---|---|---|---|---|
| 128 | 16 (1.9 M) | 410 | 365 | 64 | 45 |
| 128 | 32 (7.8 M) | 214 | 196 | 53 | 38 |
| 256 | 16 (1.9 M) | 107 | 128 | 16 | 11 |
| 256 | 32 (7.8 M) | **56** | 60 | 13 | 10 |
| 256 | RadioUNet first U (10.9 M, `--arch radiounet`) | 43 | 42 | – | – |

- Inference at 256 px, base 32: 259 maps/s (3.9 ms/map, batch 32).
- Loader (uint8 memmap reads + on-GPU featurizer): 12k samples/s cold, 20k warm. **The GPU is the only bottleneck**, so worker processes are not needed. B1 is computed on the GPU from the cached straight-line features instead of per sample on the CPU.

Decisions:

- **channels_last: dropped.** It is 4–6× slower on MPS.
- **fp16 autocast: dropped.** It gains 0–19%, and a 5-minute RMSE check isn't worth doing for such a small speed-up. All runs are fp32.
- **Backbone: our U-Net, base 32, depth 5, BatchNorm**, the same for every variant. The port of RadioUNet's own first U-Net is 25% slower, so "closer to RadioUNet" is not cheap. It is kept as `--arch radiounet` in case Phase B misses the target.

## 4. ETA

One run with the paper protocol (256 px, 50 epochs × 40,080 samples = 2.0 M samples at 56/s, plus 50 full validations of ~30 s): **≈ 10.5 h**.

| Phase | Runs | GPU hours at full protocol |
|---|---|---|
| B: geom reproduction | 1 | 10.5 |
| C: 3 variants × {50, 150, 501} maps × 3 seeds, equal steps | 27 | ≈ 280 |
| E: B2 variant, 3 seeds | 3 | ≈ 31 |

The full Phase C grid at the paper protocol is not feasible on one M2 Max. The reduced plan is in §5.

## 6. How much do the simulators disagree? (`rcm_ml/sim_agreement.py` → `runs/mps/simulator_agreement.json`)

Each simulator's map is used as the "prediction" of the other, on the full test set (99 maps × 80 Tx), with the network metrics:

| | dB (clipped) | outdoor | above −127 dB | mean bias, outdoor |
|---|---|---|---|---|
| IRT2 (ray tracing) as prediction of DPM | **3.53** | 4.01 | 4.37 | IRT2 1.0 dB weaker |

IRT2 lights 69% of outdoor pixels above the floor, against 83% for DPM (2 ray interactions miss paths that DPM routes around corners).

**Consequence:** below ~3.5 dB against DPM, a model is fitting DPM's particular behaviour rather than radio physics. The paper's 1.6 dB means "copies DPM well". A DPM-trained model's error on IRT2 should be read against this 3.5 dB floor, not against zero.

## 7. Our own ray tracer (`rcm_ml/raytrace.py`, `scripts/raytrace_eval.sh`)

A 2D shoot-and-bounce tracer on the exact building polygons (`polygon/buildings_complete`), 5.9 GHz, isotropic antennas, up to 2 interactions like IRT2:

- Path types: LOS, R, RR (7,200 rays from the Tx); D, DR, DD (fans of 1,440 rays from convex corners the source can see).
- Incoherent power sum, with a track-length estimator: free-space error is 0.00 dB mean, 0.57 dB p99.
- Not modelled: reflection-then-diffraction, ground reflection, transmission into buildings.
- Polygons are in `(x, y)` with y up. They match the raster only to IoU 0.88, i.e. about one pixel of disagreement around every building.

Test: 20 test maps × 2 Tx (the Tx that also have IRT4). RMSE in dB (clipped / outdoor / above −127 dB):

| Tracer version | vs IRT2 | vs IRT4 | vs DPM | ms/map, 1 core |
|---|---|---|---|---|
| 1. Textbook: free space, knife-edge J(ν) at corners, Fresnel concrete | 7.97 / 9.05 / 11.22 | 8.51 / 9.69 / 10.87 | 8.36 / 9.52 / 10.73 | 122 |
| 2. + distance exponent fitted on training maps (**n = 2.60**) | 5.72 / 6.48 / 8.20 | 7.23 / 8.28 / 9.32 | 6.81 / 7.79 / 8.82 | 123 |
| 3. + empirical losses, grid-searched on 12 training maps vs IRT2 (**6 dB + 20 dB per 90° at corners, 10 dB per reflection**) | **4.32 / 4.75 / 5.76** | 5.38 / 6.04 / 6.78 | 4.91 / 5.48 / 6.20 | 116 |

**WinProp is not textbook physics.** Even in line of sight, all three simulators fall off as PL = 47.6 + 26 log10 d, i.e. exponent 2.6. That is FSPL at 1 m, then 11 dB below free space by 150 m. It is smooth, so it is not a two-ray ground reflection. Its corner losses are also far gentler than the knife-edge formula: 20 dB for a 90° turn, against 40+ dB. Both look like empirical street-level settings of the simulator.

**Speed:** ~115 ms per 256 × 256 map on one CPU core (15–140 ms depending on building count). The U-Net takes ~4 ms on the GPU at 256 px. The tracer needs no training data and gets within 0.8 dB of the IRT2-vs-DPM disagreement. It is therefore a candidate physics channel for the network (a "B3"), replacing B1/B2.

## 8. Other ray tracers: Sionna RT, PyLayers, Opal

Same test maps and metrics. Sampled tracers are scored on uniform random pixels, an unbiased estimate of the full-map RMSE. Numbers are dB RMSE vs IRT2 on the clipped scale; "bias" is the tracer minus IRT2 on outdoor pixels where both are above the floor.

| Tracer | Hardware | Sample | vs IRT2 | outdoor | bias | Time | Full 256² map |
|---|---|---|---|---|---|---|---|
| Ours, calibrated (§7) | 1 CPU core | 20 maps × 2 Tx, full maps | **4.32** | 4.75 | – | 116 ms/map | **0.12 s** |
| Ours, textbook physics (§7) | 1 CPU core | same | 7.97 | 9.05 | – | 122 ms/map | 0.12 s |
| **Sionna RT 2.2**, PathSolver, 2 interactions + UTD edge diffraction, ITU concrete | Apple GPU (Mitsuba `metal_ad_mono_polarized`) | 20 maps × 2 Tx × 1,024 px | 7.35 | 8.47 | +6.7 | 8.9 s / 1,024 rx | ~9–10 min |
| **PyLayers** (2019, `d9bc71f`), signatures cutoff 2, diffraction on | 1 CPU core | 5 maps × 2 Tx × 100 px | 10.70 | 12.13 | +3.9 | 1.2 s/link | ~14 h |
| PyLayers, cutoff 3 | 1 CPU core | 1 map × 2 Tx × 50 px | 10.67 | 11.37 | +2.3 | 2.8 s/link | ~34 h |
| Opal (OptiX) | – | not runnable | – | – | – | – | needs NVIDIA GPU, OptiX 6.5, CUDA 10.2, Linux/Windows |
| *IRT2 vs DPM (two WinProp models)* | | | *3.53 vs DPM* | | | | |

Notes:

- **Sionna's RadioMapSolver cannot do this dataset.** It scores rays where they cross the measurement plane, so with Tx and receivers both at 1.5 m the map is empty (`--path-solver` is the workaround). For drones (Tx and Rx at different heights) the RadioMapSolver is the right tool and is much faster (~0.1–0.2 s/map in our smoke test). Sionna's +6.7 dB bias is the same textbook-vs-WinProp gap as our uncalibrated tracer: WinProp's d^2.6 law (§7).
- **PyLayers** needs Python 3.10 with 2022-era numpy/networkx/shapely, a stubbed basemap and a `Graph.node` shim (see `rcm_ml/pylayers_eval.py`). Its signature search is the limit: at cutoff 2, 56% of links have no path (25% at cutoff 3). Those count as "below floor", which drives both the error and the "above −127 dB" figure (19–26 dB).
- **Take-away (read the columns, not just the error):** only our calibrated tracer was tuned to WinProp (3 parameters fitted on training maps). Sionna and PyLayers ran with textbook physics, and untuned ours (7.97 dB) is on par with Sionna (7.35 dB). So the 4.3 dB reflects calibration to WinProp's empirical settings, not better physics; a Sionna run with the same tuning is still to do. The runtimes also compare different work. Ours is 2D and fills the whole map in one ray pass; Sionna's PathSolver solves 3D paths, with polarisation and phase, per pixel. Sionna's whole-map RadioMapSolver runs in ~0.1–0.2 s/map, the same range as ours, but cannot handle Tx and Rx at the same height. For 3D and drones, Sionna on Metal is the practical choice. PyLayers and Opal are not.

### 8.1 Correction: the Sionna numbers above are not reliable

Full-map renders (`rcm_ml/render_cases.py` → `figures/03_all_methods_419_0.png`, `03_all_methods_392_1.png`) showed horizontal bands in Sionna's map that follow our receiver batches. We re-ran a ~2,500-pixel patch of map 392 several ways:

- **Same pixels, row-ordered vs shuffled batches (1e6 samples):** 9.1% of pixels flip between "above floor" and "below floor".
- **1e6 vs 1e7 samples:** 22% flip, and more samples found fewer lit pixels (59% vs 74%).
- **Path cap raised to 1e7–1e8, batches of 256:** still 14–23% flips, and 1.9–2.8 dB mean difference where both runs find signal.

So in our configuration Sionna's PathSolver output depends on batching and sampling, especially for diffracted paths. The cause was not isolated. **Treat the 7.35 dB (§8) and the rendered Sionna maps as provisional: they measure our use of Sionna, not Sionna.** Before quoting a Sionna accuracy, one receiver per call (or Sionna's own guidance for dense receiver grids) should be checked.

Per-scene scores of the renders (256 px, clipped dB, `runs/mps/render/scores.json`):

| Method | map 419 vs DPM / IRT2 | map 392 vs DPM / IRT2 |
|---|---|---|
| IRT4 (WinProp) | 3.24 / 2.02 | 3.48 / 2.90 |
| B1 / B2 | 5.80 / 6.31 · 6.11 / 6.55 | 6.32 / 6.83 · 7.01 / 7.50 |
| Our tracer, textbook / calibrated | 9.82 / 10.49 · 3.57 / 3.04 | 9.53 / 9.77 · 4.12 / 3.87 |
| Sionna (provisional, see above) | 10.21 / 11.16 | 9.00 / 9.61 |
| PyLayers (8 m grid, cutoff 2) | 12.62 / 12.41 | 12.55 / 12.55 |
| U-Net geometry / + physics / + residual (128 px, upsampled) | 2.87 / 3.25 · 2.33 / 3.04 · 2.36 / 3.05 | 2.91 / 3.33 · 2.22 / 3.05 · 2.33 / 3.14 |

## 9. Bigger maps (`rcm_ml/bigmap.py` → `runs/mps/bigmap.json`, `figures/03_bigmap_1km.png`, `03_bigmap_2km.png`)

RadioMapSeer has only 256 m tiles, so we stitched k × k test cities into one map and put the Tx in the middle (seams are artificial). There is no WinProp truth at this size: our calibrated tracer (3–4 dB from IRT2 on tiles) is the **proxy reference**. U-Nets are the 128 px PoC checkpoints (2 m/px) applied unchanged to the larger input.

**Runtime per map** (tracer and B1 on 1 CPU core; U-Net on the Apple GPU while it was also training):

| Map | Wall segments | Our tracer | B1 physics map | U-Net |
|---|---|---|---|---|
| 256 m | ~200–650 | 0.12 s | 0.03 s | 0.5 ms |
| 512 m | 1,294 | 1.4 s | 0.22 s | 6 ms |
| 1 km | 5,800 | 7.3 s | 1.9 s | 15 ms |
| 2 km | 19,799 | 65 s | 15 s | 40 ms |

The tracer grows ~8× per doubling of map side, because every ray is tested against every wall (no spatial index). A uniform-grid or BVH index should bring that to ~4× per doubling. Estimate for 2 km: ~5–15 s, not measured. B1 grows the same way (one line walk per pixel). The U-Net grows with pixel count, as expected.

**Accuracy far from the Tx: false coverage**, i.e. the share of outdoor pixels the reference calls dead (below −127 dB) that a method paints as covered. Near the Tx (< 200 m) this metric is dominated by the proxy's own hard shadows, so only the far rings are meaningful:

| Method | 1 km map: 200–400 m | 400–800 m | 2 km map: 400–800 m | 800–1,500 m |
|---|---|---|---|---|
| B1 physics map | 3% | 0% | 0% | 0% |
| U-Net, geometry only | 90% | 91% | 94% | 94% |
| U-Net + physics input | 50% | 56% | 46% | 49% |

- **Networks trained on 256 m tiles do not transfer to big maps.** The geometry-only net sees ~200 px (≈ 400 m at 2 m/px) around each pixel. Beyond that it cannot tell where the transmitter is and paints faint signal (≈ −120 dB) almost everywhere. The physics input halves this, because the physics map carries distance, but half the dead area is still shown as covered.
- **At street level the signal really dies within ~200–300 m** in these cities (5.9 GHz, 1.5 m antennas), so most of a 2 km map is below the floor. For drones at altitude the covered area is much larger, and scale matters more.
- **Implications for a km-scale product:** train on larger tiles or with a larger receptive field (deeper or dilated network, multi-scale inputs), or let the physics/tracer handle the far field and use the network only near the Tx. Also add a spatial index to the tracer.

## 10. Training data for bigger maps and drone altitudes

**Existing datasets (checked 2026-09-29):**

| Dataset | Area / grid | Heights | 5.9 GHz | Simulator | Licence / availability |
|---|---|---|---|---|---|
| **FARM / ARM** (arXiv 2604.17362), `hf://datasets/jliang097/FARM_training_test` | 1 km scenes, 512 × 512 released grid; 130 OSM cities | **30 Rx layers, 5–100/150 m** | yes (iso and 30°/120° beams) | Sionna RT | MIT; 13 groups × 15k maps; ~0.6–0.8 GB per scene archive, ~80–100 GB total |
| SpectrumNet (arXiv 2408.15252) | 1.28 km, 10 m cells, 15,300 areas | Rx 1.5 / 30 / 200 m | no (0.15–22 GHz, not 5.9) | ray tracing | CC0, "to be uploaded" to github.com/ShuhangZhang/FDRadiomap (not verified) |
| UAV mmWave path loss, `hf://datasets/SHussain37/uav_pathloss_dataset` | 5 cities, 4 Tx sites | UAV Tx 25/35/45 m | no (mmWave) | ray tracing | small, CSV |
| Multi-config Radiomap (U6G XL-MIMO), `hf://datasets/lxj321/...` | 800 urban scenes | base-station | no (U6G) | – | CC BY 4.0 |

FARM matches our needs best (km scale, altitude layers, 5.9 GHz). Its transmitters are base stations placed on a grid; whether a ground-level pilot (1.5 m) is represented needs checking in the data (`tx_positions_512.csv`).

**Generating our own with Sionna RadioMapSolver** (`rcm_ml/sionna_drone.py` → `runs/mps/sionna_drone*.json`). Scene: the 1 km mosaic (25 m buildings, ITU concrete, dry-ground plane), pilot at 1.5 m, drone plane at 30/60/120 m, up to 3 interactions + edge diffraction, Apple GPU shared with training:

| Setting | Time per 1 km map | Cells with signal | Seed-to-seed |dB| mean / p95 |
|---|---|---|---|
| 2 m cells, 1e7 rays, 30 / 60 / 120 m | 1.4 s | 18% / 51% / 77% | 5.2–6.1 / 26–28 |
| 2 m cells, 1e8 rays, 60 m | 4.8 s | 87% | 6.2 / 24 |
| 4 m cells, 1e8 rays, 60 m | 5.3 s | 96% | 4.7 / 19 |

Generation is fast (thousands of km² maps per GPU-hour), but single runs are noisy in weak-signal cells, and 10× more rays barely reduces the spread. Before using Sionna maps as training targets: average several seeds, restrict the loss to cells above a signal threshold, and verify on a small scene against a converged reference.

## 11. FARM group D1 inspected (`rcm_ml/farm_inspect.py` → `runs/mps/farm_d1_inspect.json`, `figures/03_farm_scene1.png`)

Downloaded `D1` (10 scenes, 5.9 GB compressed; scene 1 extracts to 11 GB, in `data/farm/`, git-ignored).

- **Layout:** `<scene>/freq{21,33,59}/{iso,pattern_30,pattern_120}/tx<ID>_yaw<Y>.npy`, 100 Tx per scene. Each file is `uint8 (1, 30, 512, 512)`: 30 receiver altitudes (5–150 m, 5 m steps) × 512² cells.
- **Pixel size ≈ 1.95 m (1 km / 512), whole scene, not Tx-centred.** Inferred by fitting free-space loss at 120/150 m: R² = 0.991 at 1.95 m vs 0.972 at 1 m.
- **Encoding (inferred):** gain_dB = (value − 253.2) / 0.985, about 1 dB per step. Buildings = 0. The **no-signal floor is 135 (= −120 dB)**, not the 130 the README states.
- **Transmitters: all 1,000 in D1 are at 20 m** (a mast or rooftop). No transmitter at pilot height (1.5 m), and the lowest receiver layer is 5 m. Buildings are 5–20 m tall.
- **Physics is free-space-like at altitude** (slope 0.985 per dB of FSPL), unlike WinProp's empirical d^2.6 at street level.
- **Coverage above the floor, 5.9 GHz iso, scene 1:** 5 m 23%, 10 m 34%, 15 m 56%, **20 m 23%**, 25 m 99.5%, ≥ 30 m ~100%. The dip at 20 m, exactly the Tx height, matches the RadioMapSolver blind spot we found in §8 (a measurement plane at the Tx height gets no ray crossings). **Treat FARM's Tx-height layer as unreliable.**

**Fit for our drone link (pilot at ~1.5 m, drone at 30–120 m):** not a direct match. FARM models a 20 m mast talking to a drone. By reciprocity, a drone link needs one end at hand height, which FARM never has. Uses:

- pretraining on altitude-dependent, km-scale maps (the network learns 3D-ish geometry and a large receptive field)
- a benchmark for methods that must handle altitude
- a scene library: 130 OSM cities with building heights (recoverable from the layer occupancy)

For our exact geometry we still need our own data: Sionna with the pilot at 1.5 m, the drone plane well above it, several seeds averaged (§10), validated on the DJI flight logs.
