Skip to content
Source Separation Lab
Methods · decision records

Decision record DR-004 · recorded 2026-10-06

Measuring the cost of ρ = 0.625 against a comparator fixed in advance

  • Status: Accepted. It replaces the headline comparison reported under "What happened" in DR-001, which stays as written.
  • Date recorded: 2026-10-06
  • Decision in one line: On /tune the cost of ρ = 0.625 is measured against ρ = 0.60, the lowest point of the 2021 mean curve, which was fixed before any 2026 data were drawn; the comparison with the best ρ of the current run is kept only as a secondary figure, labelled as biased upwards.

Context

DR-001 reported that on 50 paired datasets ρ = 0.625 gives a higher MSE than ρ = 0.575, by 0.089 on average (95% interval 0.081 to 0.096). But 0.575 is the minimiser of the mean curve of those same 50 datasets, picked from 41 values of ρ. Choosing a comparator as the best of many candidates and then testing against it on the same data is a post-selection comparison. The winner's mean MSE is optimistically low (the winner's curse), so the gap is biased upwards, and neither the interval nor the sign test accounts for the selection.

A review of the 2026 upgrade found the problem. At the default settings it hardly matters, because every bootstrap resample puts the minimum at 0.575. At small n or other noise settings it does:

  • With temporal noise variance 0.10, spatial 0.03 and 10 paired realisations, the minimum moves to 0.60 and the gap shrinks to 0.009 (sign test p = 0.34).
  • With temporal variance 0.10, spatial 0.005 or 0.015 and 10 paired realisations, the minimum is 0.625 itself, and the site compared 0.625 with itself.

Decision

The headline is now a pre-specified paired comparison: MSE at ρ = 0.625 minus MSE at ρ = 0.60, on the same datasets. 0.60 is option 1 in DR-001, the literal grid minimum of the 2021 curve, so it was fixed by the 2021 data and not by the datasets it is tested on.

The comparison with the current run's best ρ stays as a secondary figure. It is labelled as picked on the same data and biased upwards, and it is dropped when the run's best ρ is 0.625 or 0.60. The oracle regret against each dataset's own best ρ also stays, labelled as optimistic.

Options considered

  1. Keep the in-sample best ρ as the comparator and add a caveat. This is the least work, but the headline number would still carry the bias.
  2. Compare with a pre-specified ρ = 0.60. This is what I chose.
  3. Split the sample. Choose the comparator on the first half of the realisations and test on the second half. This is valid, but it halves n, and the comparator changes as the slider moves.
  4. Use a selection-aware bootstrap. Re-select the minimiser inside every bootstrap resample, so the interval includes the selection step. The point estimate would still carry the winner's curse.

Why

Option 2 needs no adjustment and keeps the full sample. It is also the comparison a reader of the 2021 report could have made, because 0.60 was on the grid and was its minimum. It answers the question DR-001 is really about: did taking the midpoint of 0.60 and 0.65 cost anything compared with taking the grid minimum?

Options 3 and 4 answer a harder question, how far 0.625 is from the best fixed ρ, and need more machinery. The oracle regret already gives an upper bound for that.

What happened

With seed 30034 and 50 paired datasets:

  • Against the pre-specified ρ = 0.60, ρ = 0.625 adds 0.060 MSE on average (95% interval 0.057 to 0.062). It is worse on all 50 datasets (exact sign test p < 0.001), and dzd_z = 7.2. The MSE is about 10% higher than at 0.60.
  • Against the run's best ρ, 0.575, the gap is 0.089 (0.081 to 0.096), as in DR-001. It is larger, as the selection predicts, and it is now the secondary figure.
  • Against each dataset's own best ρ, an oracle that no fixed ρ can reach, the gap is 0.092 (0.084 to 0.100).

So the conclusion of DR-001 holds with a comparator chosen in advance: the midpoint rule cost reconstruction error compared with simply taking the grid minimum.

The weak points:

  • dzd_z = 7.2 looks large, but under common random numbers it measures how consistent the difference is, not how big it is. The practical size is the 10%.
  • The answer depends on the noise level. With temporal variance 0.10, spatial variance 0.005 and 10 realisations, ρ = 0.625 beats 0.60 by 0.032 (95% interval 0.013 to 0.051). At the 2021 noise level the conclusion is clear, but it does not carry over to other settings.
  • The comparison is only as meaningful as the criterion. The 2021 MSE reconstructs X from itself (DR-001), so "worse" here means a worse fit to X, not worse recovery of the sources.

What I'd change

  • Write the comparator into the analysis plan before running the simulation, not after seeing the curve.
  • To ask how far a choice is from the best fixed value, use sample splitting or a selection-aware bootstrap, not the in-sample minimum.
  • Report a practical size next to dzd_z wherever paired differences are shown, as /recovery already does.