# TASK 01: Radio-map models on the Mac GPU (for Claude Code)

**Owner:** Pavlo · **Executor:** Claude Code, running natively on macOS in `~/projects/radio-coverage`
**Status:** ready · **Written:** 2026-09-29

## 0. Why this task exists

A first pass ran on 2 cloud CPU cores: 12 minutes per model, half resolution,
under 2 epochs. It produced a promising but unproven result: a U-Net that gets
a cheap physics map as an input beats a geometry-only U-Net by ~1 dB
(3.3 vs 4.3 dB RMSE on RadioMapSeer/DPM). Converged, full-resolution,
multi-seed runs on the M2 Max GPU (PyTorch **MPS**) will show whether that holds.

Read first, in order: `docs/research/01-literature-and-datasets.md`,
`docs/research/02-radiomapseer-experiments.md`, `rcm_ml/README.md`.

## 1. Goals (in priority order)

1. **Reproduce the published baseline.** A geometry-only U-Net (RadioUNet_C-style) reaches ≤ 2.3 dB RMSE on the DPM test set at 256×256. The paper reports 0.020 gray RMSE; with the exact encoding factor that is ≈ 2.0 dB. If we can't get close, every other comparison is suspect.
2. **Answer the research question with statistics.** Does the physics input (`feats`, `hybrid`) beat `geom` at convergence? Does it hold with 50 / 150 / 500 training maps? Use 3 seeds each and report mean ± std.
3. **Test robustness to a different simulator.** Train on DPM, test zero-shot on IRT2 (ray tracing). Does the physics channel reduce the transfer gap?
4. **Build a better cheap physics channel (B2)** that follows street canyons instead of straight lines through buildings, then re-run goal 2 with it.
5. **Optional:** export the best model to Core ML and benchmark inference on the GPU vs the Neural Engine.

## 2. Non-goals

- No drone-altitude work, no Sionna, no changes to the web app (`app/`, `rcm/`). Those are the next task.
- No new model architectures beyond what's specified. The comparison must keep the backbone fixed.
- No hyperparameter search beyond what's listed. We want a fair comparison, not a leaderboard.

## 3. Hard constraints

- **Run natively, not in Docker.** Docker/VMs on macOS have no MPS. Check with `torch.backends.mps.is_available()` and fail loudly if it's false.
- **Never commit or print** `.dji_key`, `logs/`, `csv/`, `extract/`, `data/`. They are in `.gitignore`; keep it that way.
- Don't delete anything in `logs/` (the original DJI flight records).
- Every number in the docs must come from a file in `runs/` produced by committed code. Record the git commit hash in each `test.json`.
- If an assumption below turns out to be wrong (e.g. the PNG encoding), stop, document it in `docs/research/`, and fix it before continuing.

## 4. Known facts (verified; do not rediscover)

- Dataset: RadioMapSeer, Google Drive id `1_EHrS-Sp25nzmbJKbwnNVzbSR57jb6_8`, 3.26 GB zip, CC BY 4.0. Needed folders: `png/buildings_complete`, `png/antennas`, `gain/DPM`, `gain/IRT2`, `antenna`, `dataset.csv`.
- Encoding: gray `f ∈ [0,1]`, `PL_dB = −147 + 99.16·f`. The paper says "dB RMSE = 80 × gray RMSE"; the encoding gives 99.16. Report **gray RMSE** (comparable to the paper) and **dB RMSE = gray × 99.16**, plus RMSE restricted to pixels above −127 dB.
- Antenna PNG row = 255 − json y. Tx is always inside the central 150×150.
- DPM targets can be non-zero inside buildings (seen up to gray 156). The paper evaluates all pixels, so we do too. Also report RMSE on outdoor pixels only.
- Split: `rcm_ml.data.split(seed=0)` is a seeded 500/100/100 over map ids 0–699. It is **not** the paper's exact split. Check whether the official split file is on radiomapseer.github.io or github.com/RonLevie/RadioUNet. If it is, add it as `split(official=True)` and use it for goal 1.
- Our CPU results to beat (128 px, 12 min): geom 4.31, feats 3.29, hybrid 3.42, hybrid-50-maps 3.37 dB; B1 alone 15.0 dB (4.8 above −127).

## 5. Work plan

### Phase A: setup and sanity (≤ 1 h)
0. The folder is not a git repo yet. `git init`, check that `.gitignore` covers the private paths in §3, and make an initial commit.
1. Run `scripts/run_mac.sh` only up to the 5-minute benchmark (comment out the final loop), or run the steps by hand.
2. `pytest -q` passes. Add tests for `rcm_ml`: encode/decode round trip, Tx pixel location, baseline fit reproduces slope ≈ −22 dB/decade.
3. Measure throughput (samples/s) on MPS at 128 and 256 px, batch 15, for `base=16` and `base=32`. Record it in `docs/research/03-mps-runs.md`.
4. Try `channels_last` and `torch.autocast("mps", dtype=torch.float16)`. Keep them only if RMSE on a 5-minute run doesn't get worse.

**Done when:** throughput is known, and there's an ETA for Phase B written in the doc.

### Phase B: goal 1, reproduce RadioUNet (geom)
- 256 px, all 80 Tx per training map (`prepare --tx-train 80 --tx-eval 80`), Adam 1e-4, batch 15, 50 epochs, best-on-val checkpoint (the paper's protocol). Backbone `base=32`, or closer to RadioUNet if it's cheap to match.
- The training loop needs **epoch-level validation on the full val set** (not 8 batches) and a `--seed` flag. Add both.

**Done when:** test RMSE for geom is reported. If it's > 2.3 dB, investigate (split, resolution, capacity, epochs) before moving on, and write down what you found.

### Phase C: goal 2, physics input vs geometry at convergence
- Variants: `geom`, `feats`, `hybrid` × train maps {50, 150, 500} × seeds {0, 1, 2}. That's 27 runs; start with 500 maps × 3 seeds × 3 variants.
- Same epochs for all. With fewer maps, use more epochs so the number of optimizer steps is comparable, and state which choice you made.
- Report a table of mean ± std for gray RMSE, dB RMSE, dB RMSE above −127 dB, and outdoor-only dB RMSE. Plot error vs training maps (log x).

**Done when:** `docs/research/03-mps-runs.md` has the table and plot, and states which differences are larger than seed noise.

### Phase D: goal 3, DPM → IRT2 zero-shot
- Evaluate the Phase C checkpoints (500 maps) on the IRT2 test targets. Add `--eval-target irt2` to `train.py` or a small eval script.
- Also report the IRT2-trained geom as an upper reference.

**Done when:** there's a table of the transfer gap per variant.

### Phase E: goal 4, street-canyon physics channel B2
Cheap, deterministic, no learning:
- Geodesic distance from Tx through free space only (8-connected Dijkstra or fast marching on the 256² building mask). Count "turns": where the geodesic path direction changes by more than 30°, or use the parent-pointer bend count.
- B2 model: `PL = a + b·log10(geodesic_d) + c·turns + e·LOS + g·(straight-line metres inside buildings)`, fitted like B1 (least squares on above-floor training pixels).
- Target runtime < 50 ms per map on CPU (use numba or scipy's `dijkstra` on the grid graph).
- Report B2 alone vs B1 alone. Then re-run the best variant from Phase C with B2 instead of B1 (500 maps, 3 seeds).

**Done when:** B2's accuracy and runtime are reported and the network comparison is done.

### Phase F (optional): Core ML / Neural Engine
- Export the best model with `coremltools` and measure latency on `CPU_AND_GPU` vs `CPU_AND_NE` compute units, plus the accuracy delta after fp16. Record in the doc.

## 6. Deliverables

- Code changes committed in small commits (one per phase), with the messages the repo uses.
- `docs/research/03-mps-runs.md`: setup, throughput, all tables, plots in `docs/research/figures/`, and an honest "what we can and can't conclude" section.
- An updated "Next" section in `docs/research/02-radiomapseer-experiments.md` pointing at 03.
- `runs/**/test.json`, each with the commit hash, seed, config, and wall time.

## 7. Report back to Pavlo (at the end, ≤ 15 lines)

1. Did we reproduce RadioUNet (number vs 2.0 dB)?
2. Physics input vs geometry at convergence: effect size ± noise, per data size.
3. DPM → IRT2 transfer gap per variant.
4. B2 vs B1.
5. Anything that contradicts `02-radiomapseer-experiments.md`.
6. Total GPU hours used.

## 8. Risks and gotchas

- **MPS gaps:** some ops fall back to CPU silently. Set `PYTORCH_ENABLE_MPS_FALLBACK=1` only if needed, and log a warning when it happens.
- **Memory:** the full cache is ~12 GB of uint8 memmaps (40k samples × 4 arrays × 64 KB). Fine on 64 GB, but keep memmaps and don't load everything into RAM.
- **Data loading may bottleneck** once the GPU is fast. `num_workers=0` is set now because of memmaps. Try 4 workers with `persistent_workers=True`, and precompute the B1/B2 maps into the cache instead of recomputing them per sample.
- **Seed noise may be as large as the effects.** That is a valid result; report it rather than cherry-picking.
