Skip to content
Source Separation Lab
Methods · decision records

Decision record DR-005 · recorded 2026-10-06

Student-t intervals for means over fewer than 30 realisations

  • Status: Accepted
  • Date recorded: 2026-10-06
  • Decision in one line: Means over 30 or more Monte-Carlo realisations keep the 95% percentile-bootstrap interval, and means over fewer than 30 use the Student-t interval, because the percentile bootstrap undercovers at small n.

Context

Every 2026 page first put a 95% percentile-bootstrap interval on every mean over realisations: 2,000 resamples, with indices from numpy's generator so that Python can reproduce each interval. Several controls allow small samples. The noise-sensitivity heatmap on /recovery defaults to 10 realisations per cell, and /tune and /bias-variance go down to 10.

For a mean, the percentile bootstrap behaves like a normal interval built on the plug-in standard deviation (ddof = 0). It is too narrow when n is small. I ran a simulation with normal data and 4,000 runs per sample size. A nominal 95% percentile interval covered the true mean in 90% of runs at n = 10, 92% at n = 20, and 94% at n = 30, 50 and 100. A review of the upgrade had flagged this, so the "95%" labels overstated the coverage at the default heatmap setting.

Decision

Add meanCI and meanCurveCI to web/src/lib/stats/.

  • Fewer than 30 values: the Student-t interval xˉ±tn−1, 0.975 s/n\bar x \pm t_{n-1,\,0.975}\, s / \sqrt n.
  • 30 values or more: the percentile bootstrap, unchanged.

Every mean on the site now goes through these functions. The distribution of the curve's minimiser on /tune stays a bootstrap at every n, because it is not a mean.

Each estimate records which method it used. The figure captions and the AI figure summaries say "Student-t" when it applies.

Options considered

  1. Keep the bootstrap and label small-n intervals as approximate. This is honest, but the default heatmap would still show intervals that are known to be too narrow.
  2. Raise every minimum to 30 realisations. The heatmap would take three times as long, and /tune could no longer run the 2021 ten on their own.
  3. Use the Student-t interval below 30. This is what I chose.
  4. Use bootstrap-t, BCa or the expanded percentile bootstrap at every n. These are better in theory, but they would change every published number at n ≥ 30 and need more machinery to test against scipy.

Why

The quantities averaged here are smooth and light-tailed: MSEs, correlations and their paired differences. The t interval is exact for normal data and close for these. In the same simulation it covered the true mean in 94% to 95% of runs at n = 10 and 20.

The switch is also cheap to verify. The new t quantile is checked against scipy.stats.t.ppf to a relative 1e-10 for 1 to 1,000 degrees of freedom. The interval matches scipy.stats.t.interval.

Keeping the bootstrap from 30 realisations up leaves every default result unchanged (50 or 100 realisations), and those results stay reproducible in Python.

What happened

  • Every number quoted at the default settings is unchanged.
  • At 10 paired realisations on /tune, the interval for the cost of ρ = 0.625 against 0.60 is now 0.055 to 0.064 (Student-t) around a mean of 0.060.
  • The heatmap's per-cell intervals are wider at its default of 10 realisations per cell.
  • The weak point is the threshold itself. From 30 realisations up, the percentile bootstrap still covers about 94% rather than 95% for a mean of normal data. The 30 is a convention, not a cliff.

What I'd change

  • Use one interval method at every n. The expanded percentile bootstrap (Hesterberg 2015) is the obvious candidate, because it keeps the bootstrap and fixes the small-n narrowness. Adopting it later would need a new record, because it changes the published numbers.
  • Choose the default number of realisations from the precision the reader needs, not from how long the page takes to compute.