Skip to content
Source Separation Lab
Methods · decision records

Decision record DR-002 · recorded 2026-10-06

Which noise to reproduce for parity: numpy's stream, with the R-generated noise kept as a check

  • Status: Accepted
  • Date recorded: 2026-10-06
  • Decision in one line: The browser port regenerates the noise the 2021 Python notebook actually used, by porting numpy's legacy random generator bit for bit, and keeps the R-generated noise in coursework/data/ as provenance and as a parity check, not as the source of the reported numbers.

Context

The coursework folder contains two sets of noise for the same assignment.

  • coursework/data/gamma_t.csv and gamma_s.csv were drawn in R with set.seed(30034) by a small R notebook in the same folder. Nothing in the final Python notebook reads them.
  • The final Python notebook draws its own noise with np.random.seed(30034) and np.random.normal. A comment in that cell notes that the Python generator led to a different result. Every staged output (X, the standardised X, the 10 × 21 MSE table and the correlation vectors) comes from this numpy noise.

When I revived the project as a website in 2026, every figure had to be recomputed in the browser and match the report. That meant deciding which noise the browser should reproduce, and how. One option on the table was to reuse the R-generated files, because they are already in the repository.

Decision

Port numpy's legacy RandomState (the MT19937 generator and its polar Gaussian sampler) to TypeScript, regenerate the noise from the seed, and treat the numpy noise as the reference. The R files stay where they are. A test recomputes the 2021 pipeline on them and checks the result against an independent numpy run, so the R noise is documented rather than discarded.

Options considered

  1. Reuse the R-generated noise files. They are small, about 4,000 numbers, but the report was not computed from them, so every reported number would differ.
  2. Ship numpy's noise as data. Q1 needs about 4,000 numbers, but the Q2.3 Monte-Carlo used 210 fresh draws of about 4,000 numbers each, close to 860,000 numbers. That is several megabytes, and it would only ever reproduce seed 30034.
  3. Port the generator. No data shipped, every seed works, and the extra realisations in the 2026 analyses continue the same stream. The cost is a careful port with tests.
  4. Use any JavaScript generator and accept different numbers. This loses parity with the report, which is the main reason the site exists.

Why

Only option 3 reproduces every staged output, including all 210 Monte-Carlo fits, without shipping data. It also lets visitors change the seed and keeps the 2026 extensions reproducible. Python users can recreate any realisation with np.random.RandomState(seed).

What happened

  • The parity tests pass. The standardised X matches the staged CSV to 1e-11, and all 210 entries of the MSE table match to a relative 1e-9.
  • The same port later gained numpy's randint, which drives the bootstrap. Its output matches RandomState(seed).randint exactly, so every bootstrap interval on the site can be reproduced in Python (scripts/export_stats_reference.py does this).
  • Running the 2021 pipeline on the R noise gives ΣcT,LSR\Sigma c_{T,\mathrm{LSR}} = 5.2510, ΣcT,RR\Sigma c_{T,\mathrm{RR}} = 5.2552 and ΣcT,LR\Sigma c_{T,\mathrm{LR}} = 5.4476, against 5.2674, 5.2713 and 5.4275 in the report. With the as-submitted map scoring, ΣcS,RR\Sigma c_{S,\mathrm{RR}} = 1.1903 and ΣcS,LR\Sigma c_{S,\mathrm{LR}} = 1.8301, against 1.1487 and 1.8078. The ordering of the methods is the same on both, and the numbers differ in the second or third decimal place. The R noise was a valid realisation that simply would not reproduce the report.

The weak point is that the port depends on numpy's legacy generator. numpy has frozen RandomState for backward compatibility, so this is stable, but the newer Generator API (PCG64) is not supported.

What I'd change

  • In 2021, I would have saved the exact noise the notebook used next to the outputs, or recorded the generator and seed in the outputs themselves.
  • I would generate data in one language only. Two seeded generators with the same seed in two languages produce two different datasets, and the commented-out code made it unclear which one the results came from.