- Status: Accepted
- Date recorded: 2026-10-06
- Decision in one line: The browser port regenerates the noise the 2021 Python notebook actually used, by porting numpy's legacy random generator bit for bit, and keeps the R-generated noise in
coursework/data/as provenance and as a parity check, not as the source of the reported numbers.
Context
The coursework folder contains two sets of noise for the same assignment.
coursework/data/gamma_t.csvandgamma_s.csvwere drawn in R withset.seed(30034)by a small R notebook in the same folder. Nothing in the final Python notebook reads them.- The final Python notebook draws its own noise with
np.random.seed(30034)andnp.random.normal. A comment in that cell notes that the Python generator led to a different result. Every staged output (X, the standardised X, the 10 × 21 MSE table and the correlation vectors) comes from this numpy noise.
When I revived the project as a website in 2026, every figure had to be recomputed in the browser and match the report. That meant deciding which noise the browser should reproduce, and how. One option on the table was to reuse the R-generated files, because they are already in the repository.
Decision
Port numpy's legacy RandomState (the MT19937 generator and its polar Gaussian sampler) to TypeScript, regenerate the noise from the seed, and treat the numpy noise as the reference. The R files stay where they are. A test recomputes the 2021 pipeline on them and checks the result against an independent numpy run, so the R noise is documented rather than discarded.
Options considered
- Reuse the R-generated noise files. They are small, about 4,000 numbers, but the report was not computed from them, so every reported number would differ.
- Ship numpy's noise as data. Q1 needs about 4,000 numbers, but the Q2.3 Monte-Carlo used 210 fresh draws of about 4,000 numbers each, close to 860,000 numbers. That is several megabytes, and it would only ever reproduce seed 30034.
- Port the generator. No data shipped, every seed works, and the extra realisations in the 2026 analyses continue the same stream. The cost is a careful port with tests.
- Use any JavaScript generator and accept different numbers. This loses parity with the report, which is the main reason the site exists.
Why
Only option 3 reproduces every staged output, including all 210 Monte-Carlo fits, without shipping data. It also lets visitors change the seed and keeps the 2026 extensions reproducible. Python users can recreate any realisation with np.random.RandomState(seed).
What happened
- The parity tests pass. The standardised X matches the staged CSV to 1e-11, and all 210 entries of the MSE table match to a relative 1e-9.
- The same port later gained numpy's
randint, which drives the bootstrap. Its output matchesRandomState(seed).randintexactly, so every bootstrap interval on the site can be reproduced in Python (scripts/export_stats_reference.pydoes this). - Running the 2021 pipeline on the R noise gives = 5.2510, = 5.2552 and = 5.4476, against 5.2674, 5.2713 and 5.4275 in the report. With the as-submitted map scoring, = 1.1903 and = 1.8301, against 1.1487 and 1.8078. The ordering of the methods is the same on both, and the numbers differ in the second or third decimal place. The R noise was a valid realisation that simply would not reproduce the report.
The weak point is that the port depends on numpy's legacy generator. numpy has frozen RandomState for backward compatibility, so this is stable, but the newer Generator API (PCG64) is not supported.
What I'd change
- In 2021, I would have saved the exact noise the notebook used next to the outputs, or recorded the generator and seed in the outputs themselves.
- I would generate data in one language only. Two seeded generators with the same seed in two languages produce two different datasets, and the commented-out code made it unclear which one the results came from.