First bounded experiment — reference sensitivity audit

A no-training experiment proposal to test reference dependence before recalibration or model expansion.

Experiment ProposalDraft record

Decision and hypothesis

Recommended next experiment: reference sensitivity audit on existing RadioMapSeer assets. Hypothesis: reference choice and interaction depth change coverage/site-ranking conclusions enough that further model training should wait for an explicit reference policy. Test the documented IRT2/IRT4 discrepancy on common held-out cases, with multiple thresholds and positive-class metrics.1

This is a proposal for a later scope and budget approval. It has not been run here. It does not require new training, paid APIs or downloads. The present assessment authorizes its design only.

Inputs and prerequisite stop gate

Use existing local DPM, IRT2 and IRT4 map assets, scene/building masks, official map split and transmitter coordinates, plus the current tracer/code configuration. Record code commit, data hashes, licensing/access context, exact decoding and floor. Preserve original inputs. Check checkpoint availability only if an optional existing-model comparison is separately included; no such availability is assumed.1

Stop before execution if common reference maps or reliable decoding/Tx coordinates are missing. Produce an availability/error report instead of substituting generated labels. Do not fetch data or treat near/far mosaics as independent physical truth. Existing encoding corrections in the project note supersede the original task recipe.1

Fixed design

  1. From each official train/validation/test map partition, choose 20 available common maps with a fixed random seed (proposed 20261009). Sample eight common Tx per map without replacement. Save IDs before scoring. This yields at most 480 scene–Tx cases. Report exclusions, city mix and availability bias; all Tx from a map stay in the same split.
  2. Compute DPM/IRT2/IRT4 pairwise differences from existing maps on an identical outdoor mask/resolution. Use the declared path-gain sign convention consistently; proposed threshold grid is −110, −120 and −127.2 dB path gain, chosen from existing studies to expose threshold dependence. These are analytical thresholds, not detector sensitivities. Compute coverage using raw decoded reference arrays with strict gain > threshold before any training-floor clipping. Keep original no-ray/encoding-censor flags separately. The loader floor is exactly −147 + 0.2 × 99.16 = −127.168 dB; the rounded −127.2 threshold is lower and would misclassify floor-clipped dead cells as covered. Preserve unrounded floor values in computation. If only clipped arrays are available, coverage at or below that floor is unidentifiable for censored cells: report unknown/bounds and stop the affected comparison instead of counting them as covered.
  3. On the 160 test cases only, compare the existing tracer at two versus four interactions with otherwise fixed current calibration. No fitting or hyperparameter search in this first audit. If runtime is excessive, stop and report the completed predeclared subset, without substituting a favorable set. Any recalibration is a separate experiment using train/validation only.
  4. Compare one-site ranking regret within each test map's eight candidate sites, computed relative to that reference's maximum coverage. This is not directly comparable with the historical 80-site regret. Report rank ties and all pairwise reference changes. Do not claim a global multi-site optimum from a greedy reference.
  5. Score final test once after fixing masks/decoding/metrics. Inspect train/validation only for pipeline checks; if the test is used diagnostically, label it exploratory and reserve new maps for subsequent qualification.

Metrics and uncertainty

All comparisons are simulator/reference comparisons, not measured RF errors or detector recall.

Proposed decision and stop criteria

Before execution, register a reference-sensitive flag if the absolute pairwise coverage-share difference exceeds 5 percentage points with a scene-bootstrap interval excluding zero in a supported ring/threshold, or mean site-choice regret against another reference exceeds 1 pp. These are proposed tolerances for this audit, not user-approved operational requirements. If results are below those cutoffs, report bounded reference agreement; do not conclude physical correctness. Sparse rings and wide intervals produce an inconclusive result.

If sensitivity is material, choose/reference the intended simulator task and authorize a separate calibration comparison before training. If not material within this design, compare the existing baseline/hybrid under the agreed reference and target utility. For claimed detector performance, field evidence is still required. The audit cannot validate kilometer-scale extrapolation, drone-altitude coverage or the actual receiver contract.

Resource limits and reproducibility

Proposed hard execution caps for later approval: one local CPU process, at most two hours wall clock, no training, no GPU requirement, no network/paid calls. Reference-map arithmetic runs first; tracer smoke test uses the first 16 predeclared test cases, with a proposed 20-minute cap. Estimate remaining runtime from that smoke test. If projected runtime exceeds the remaining cap, stop before full tracer evaluation and request a revised budget/subsample. This cap is a proposed budget, not a runtime estimate or current authorization; available machine/data constraints are not yet measured.

New outputs: immutable input/config manifest with IDs/hashes and code commit; per-case metric table; grouped metrics/intervals; timing report and a concise decision note. Save any proposed experiment code in the source project only under its later implementation scope; the assessment repository contains no new runner. Preserve existing untracked arrays/checkpoints.

Why this experiment first

The recorded long-range/calibration disagreements weaken interpretation of every learned-map branch. Learned-ray map/speed validation remains a good secondary experiment if acceleration is the limiting requirement, but it cannot settle whether its teacher predicts reality. More SpectrumNet training does not itself resolve reference mismatch, encoding or coarse-grid transfer.12


  1. Main findings cite the IRT2/IRT4 ring discrepancy, split, encoding corrections and ranking/convergence protocols. ↩↩↩↩

  2. Learned-ray teacher/map gap and SpectrumNet transfer/seed sensitivity, supported by archived code/results. ↩

Sources, provenance and record details
Record type
Experiment Proposal
Status
draft
Generated
by: codex/gpt-6.1-sol at: '2026-10-10T05:37:12.974842Z'
Recorded checks
No verification metadata recorded.

Sources

All record metadata
type: Experiment Proposal
title: First bounded experiment — reference sensitivity audit
description: A no-training experiment proposal to test reference dependence before
  recalibration or model expansion.
status: draft
generated:
  by: codex/gpt-6.1-sol
  at: '2026-10-10T05:37:12.974842Z'
sources:
- id: main
  resource: /docs/project-information/ideas/IDEA-002-ai-signal-coverage/data/main-findings.md
  title: Main implementation/results evidence
- id: work
  resource: /docs/project-information/ideas/IDEA-002-ai-signal-coverage/data/worktree-findings.md
  title: Worktree implementation/results evidence