# Experiment 07: neural networks that learn ray tracing

_Started 2026-09-30. Code: `rcm_ml/raypaths.py` (teacher trajectories), `rcm_ml/raynet.py` (models). Raw results: `runs/raynet/**/test.json` (each has the commit hash and config)._

## 0. Plan

Grow the model in small steps, each scored against our own 2D tracer (`rcm_ml/raytrace.py`, §7 of `03-mps-runs.md`):

1. One ray: scene + Tx + launch angle → the ray's polyline (hit points, interaction types, power) up to N bounces.
2. 100 rays per Tx, batched, sharing the scene encoding.
3. Aggregate predicted rays into a radio map with the tracer's deposit; score against the tracer and WinProp IRT2; time it.
4. Next steps: more rays, diffraction fans, learned power, big mosaics.

## 1. Stage 1: one ray

### 1.1 The teacher: per-ray trajectories

`raytrace.trace` only adds power into pixels. `raypaths.ray_path` follows one ray from the Tx exactly like `raytrace._shoot` and records its polyline:

- every hit point
- how each leg ends: REFLECT, EXIT (leaves the 256 m tile), or STOP (the wall after the last allowed bounce)
- the wall index and the incidence cosine
- the power carried into each leg

`raypaths.deposit_paths` turns polylines back into a power map with the tracer's own `_deposit`. `tests/test_raypaths.py` checks it reproduces the LOS/R/RR part of `trace` to 0.001 dB on 99.9% of pixels, on a synthetic box and on map 1. So the polyline is a complete description of what the tracer does with a reflected ray. Diffraction fans are left for later.

Data (`python -m rcm_ml.raynet data`, `data/raynet/`, git-ignored; sizes and build time in `runs/raynet/data.json`):

| split | maps | Tx per map | rays per Tx | rays |
|---|---|---|---|---|
| train | 501 (official split) | 16 random of 80 | 32 random angles | 256,512 |
| val | 100 | 2 random | 32 | 6,400 |
| test | 99 | Tx 0 and 1 | 64 | 12,672 |

Rays are traced up to 4 reflections. Evaluation uses 2, like the tracer: same legs, with the 3rd hit marked STOP.

Scene input: the building polygons rasterised at 0.5 m with 4 × 4 supersampling, so pixels hold the covered fraction (anti-aliased). The polygons are the tracer's geometry; the dataset PNG matches them only to IoU 0.88 (§7 of 03). Check (`runs/raynet/data.json`): sampled at the true hit points, coverage is 0.56 on average, and 0.12 / 0.92 half a metre before / after. The raster is aligned with the walls to within a few centimetres.

### 1.2 Representation: why autoregressive next-hit in the ray's frame

A trajectory is a chain of identical sub-problems: from a point, in a direction, find the first wall and its orientation, then mirror. Two ways to learn it:

- **direct**: one CNN pass over the whole tile (occupancy, Tx, launch angle, coordinates as channels) regresses all hit points and leg types at once. This is the "image in, polyline out" model. It has to learn ray–segment intersection globally and implicitly, for every angle. The target is discontinuous in the angle: a ray grazing a corner jumps to a different wall. Errors at the first hit also change everything after it.
- **step** (chosen): an autoregressive next-hit model. For the current leg (origin o, direction d), the occupancy is sampled on a strip in the ray's own frame:
  - forward: 0.5 m bins from −2 m to the tile edge (736 bins)
  - across: 13 lanes, ±3 m

  A small 1D CNN (lanes as channels, residual dilated convolutions) outputs, for every bin i, a hit hazard h_i, a sub-bin offset and the wall normal in the ray frame. Then:
  - P(first hit in bin i) = σ(h_i) · Π_{j<i} (1 − σ(h_j)). This is the transmittance of volume rendering.
  - P(leaves the tile) = survival through all bins before the edge.

  Loss: NLL of the true bin (or of leaving), L1 on the offset, and 1 − cos on the normal at the true bin. The next direction is d mirrored about the predicted normal.

Why the step model:

- **Invariance by construction.** Translation and rotation equivariant. Map size doesn't matter, because the cost per leg is fixed, which is what bigger mosaics need.
- **"First" is built in.** The model only has to recognise a wall surface locally. The transmittance product picks the first one, and a far wall cannot win over a near one.
- **The ray stays explicit.** Each step outputs a hit point and a normal. Interaction types, power (from distance and incidence angle) and the deposit stay exact and reusable (stage 3).
- **Stated honestly:** the ray marching (where to look) is hand-coded. The network learns what a wall is, where exactly the ray meets it, and which way it faces. The direct model is the control without this prior.

### 1.3 Results

Training: 25 min wall clock per model on the M2 Max GPU, shared with a Sionna job running in another session. AdamW with a cosine schedule. The checkpoint kept is the best on validation, scored as rollout exact match at 1 m.

The step models are trained on every leg of the 4-reflection training rays (917,344 legs, teacher-forced). The direct model is trained on the 2-reflection rays (256,512).

Test set: 12,672 rays on 99 held-out maps, up to 2 reflections, so up to 3 legs per ray.

- **Exact match**: every leg has the right type (REFLECT / EXIT / STOP) and every end point is within τ of the tracer's. A ray that ends early must also end in the model.
- **Rollout**: the model feeds on its own previous hit and direction.
- **Teacher-forced**: every leg starts from the tracer's own origin and direction, which measures one-step accuracy.

| Model (run) | Params | Exact match @ 0.5 / 1 / 2 m, rollout | Same, teacher-forced | Leg-1 hit error, median / p90 | Leg-2 / leg-3 within 1 m, rollout | Wall-normal error, median / p90 |
|---|---|---|---|---|---|---|
| direct CNN (`s1_direct`) | 7.9 M | 0.4% / 1.3% / 3.9% | – | 3.57 m / 16.6 m | 0.4% / 0.3% | – |
| step, strip normal (`s1_step`) | 152 k | 28.3% / 38.7% / 52.1% | 96.2% / 98.0% / 98.6% | 0.05 m / 0.18 m | 48.8% / 26.7% | 0.69° / 2.38° |
| **step + patch normal** (`s1_step_patch`) | 1.34 M | **67.0% / 79.6% / 87.2%** | 96.0% / 97.9% / 98.5% | 0.05 m / 0.18 m | 87.9% / 75.0% | **0.06°** / 0.41° |

Longer rollouts, step + patch, up to 4 reflections (`test_rollout_4_bounces`): exact match 63.4% at 1 m. Leg 4 is within 1 m for 51.6% of rays and has the right type for 93.4%.

Inference: the 12,672 test rays (≤ 3 legs each) take 3.6 s on the GPU, including the transfer of every leg back to the CPU. It is not optimised and not yet compared with the tracer; that comparison is stage 3.

Figure: `figures/07_rays.png`. It shows 20 test rays from the first Tx of three held-out maps, as traced by the tracer, the step + patch model and the direct model.

![rays](figures/07_rays.png)

### 1.4 What this shows

1. **The representation decides almost everything.** The same data and the same 25 minutes give two very different results:
   - The direct "image in, polyline out" CNN does not learn ray tracing. Its first hit is off by 3.6 m (median), it gets 1% of rays right, and its later legs are free-hand scribbles (figure, right column).
   - The step model finds the first wall to 5 cm (median) and 0.18 m (p90), with the right leg type 99.9% of the time.
2. **One step is solved; error growth along the ray is the problem.** Teacher-forced, the step model matches 98% of rays to within 1 m. In rollout it loses rays at every bounce. A wall-normal error δ turns the reflected ray by 2δ, and that grows with distance: 1° over a 40 m leg is 1.4 m. Near corners and gaps, it also sends the ray to another wall.
3. **Precise normals give the biggest gain.** Two changes together took the median normal error from 0.69° to 0.06°, and rollout exact match from 39% to 80%:
   - An 8 m patch around the hit, sampled at 0.25 m, gets its own small CNN.
   - The normal loss is L1 on the vector instead of 1 − cos, which gives a stronger gradient near the optimum. This change alone took the strip head's own normal error from 0.69° to 0.17°, so the two changes are confounded.

4. **Still under-trained.** Validation exact match was still rising (two small dips) and reached 80.2% at the last evaluation, at 25 min (`log` in the JSON). More time, more Tx per map, or jittered origins during training (so the model sees its own errors) should help.
5. **Exact match is a harsh metric for radio maps.** A ray that grazes a corner on the "wrong" side follows a different but equally plausible path. The tracer itself is that sensitive to the launch angle, and with 7,200 rays the map averages over it. Whether 80% exact rays make a good map is the stage 3 question.

Caveats: one seed per model, 25-minute budgets, and a GPU shared with another job (so the timings are rough). Only specular reflection is covered: no diffraction yet. Power is not learned; with the calibrated 10 dB per reflection it follows exactly from the leg lengths and count.

### 1.5 Proposal for stage 2 (+ the start of 3)

The step model already runs batched, and the scene encoding (the raster) is already shared by every ray of a Tx. So "100 rays per Tx" is mostly about **evaluation**, not new machinery:

1. **Fixed fans.** For each test Tx, trace 100 rays on a uniform angle grid (the tracer's layout), with the step + patch model and with the teacher. Report:
   - exact match per Tx
   - the fraction of each Tx's fan that is right
   - throughput: rays/s on the GPU vs the numba tracer on one core
2. **Deposit straight away** (the first half of stage 3). Run predicted and true fans through `raypaths.deposit_paths` and compare the LOS/R/RR power maps: outdoor RMSE in dB, floor −127 dB. Do it at 100 rays and at the tracer's 7,200, because the track-length estimator needs dense fans. This shows whether 80% exact rays already give a near-identical map, which is my expectation (point 5 above).
3. **Cheap accuracy fixes before scaling:**
   - train ~1 h
   - add origin/direction jitter during training, so the model learns to recover from its own errors
   - optionally, a vector-scene variant: attention over nearby wall segments gives exact normals by construction, at the cost of map-size-dependent input

Diffraction (D, DR, DD fans) is the big missing piece for matching the full tracer and IRT2. It goes into stage 4: a corner-detection head, plus fans launched from predicted corners.
