MAST30034 Applied Data Science · University of Melbourne · 2021
Source Separation Lab
Six hidden signals, mixed together and buried in noise. This is my 2021 applied data science assignment, rebuilt so you can run it: generate the data, then recover the sources with least squares, ridge, lasso and principal component regression, live in your browser.
Every number here is recomputed in the browser by a TypeScript port of the original Python, and tested against the outputs saved by the 2021 notebook.
TC · 240 × 6
SM · 6 × 441
X · 240 time points × 441 pixels
The coursework
What the assignment asked
Functional MRI records how activity changes over time at thousands of points in the brain. A common simplification is that the recording is a sum of a few sources, each with a time course (when it is active) and a spatial map (where it is active).
The assignment builds a small version of this problem with known answers: six boxcar time courses, six square patches on a 21 × 21 grid, and Gaussian noise. The task is to recover the maps and time courses by regression, choose penalties sensibly, and explain why some estimators do better than others.
What I built
In 2021, and now
The original is a single Jupyter notebook in Python (numpy, pandas, scikit-learn and seaborn), with a written report. It generates and stages the data, fits each estimator, and runs a ten-realisation Monte-Carlo to pick the lasso penalty.
This site ports those algorithms to TypeScript, including numpy's seeded random generator. The defaults reproduce the 2021 matrices exactly, and every figure updates as you change the inputs. The quirks of the original code are kept on purpose and documented, so the numbers match the report.
Key results from the report
What the numbers said
- Time-course recovery, ΣcT
- 5.428
- Lasso beats ridge (5.271), which beats least squares (5.267), out of a maximum of 6.
- Spatial-map recovery, ΣcS
- 1.15 → 1.81
- Ridge to lasso. The ℓ₁ penalty clears false positives from the background of every map.
- Best lasso penalty
- ρ = 0.60
- Lowest mean MSE (0.619) over 10 noise realisations. The report used ρ = 0.625.
- Weakest principal component
- σ6 = 6.05
- PCR on all six components retrieves the sources worst of the four methods.
Explore
Four steps, one dataset
- 01GenerateBuild six time courses and six spatial maps, add noise, mix them into X.
- 02RetrieveRecover the sources with least squares, ridge, lasso and PCR.
- 03TunePick the lasso penalty ρ from the mean MSE over noisy realisations, then see how sure that choice is.
- 04Principal componentsSee why regressing on the PCs of the TCs loses their shape.
Added in 2026
Beyond the assignment: how sure are the answers?
The 2021 report answered each question with one noisy dataset. The revival keeps those answers exactly and adds the uncertainty around them. Every new result states its sample size and seed, means and proportions carry 95% intervals, and every comparison between methods is paired on the same data.
Some of the original reasoning holds up and some does not. The lasso beats ridge in all 100 simulated datasets (the 2021 one and 99 fresh draws), while the chosen penalty ρ = 0.625 turns out to sit on the steep side of the error curve.
About this project
MAST30034 Applied Data Science, Assignment 1
- Subject
- MAST30034 Applied Data Science
- University
- University of Melbourne
- Taken
- Semester 2, 2021
- Format
- Individual assignment: notebook and written report
- Author
- Sunchuangyu (Rin) Huang
- Source
- GitHub repository (private for now)
Original stack (2021)
- Python 3.9 in a Jupyter notebook
- numpy, pandas, scipy, scikit-learn
- matplotlib and seaborn for figures
- R, used once to draw a noise sample
- Report written in LaTeX on Overleaf
Revived stack
- Next.js App Router, React, TypeScript
- Framework-free TypeScript ports of every algorithm
- numpy's MT19937 generator, bit for bit
- A Web Worker for the Monte-Carlo
- Hand-built SVG and canvas charts, KaTeX, Tailwind CSS
- Vitest parity tests against the saved 2021 outputs
- Statistics checked against numpy, scipy and scikit-learn
- Optional bring-your-own-key AI, audit-logged locally
Academic integrity. The original notebook and report are kept in the repository's coursework/ folder as a record of the work. The assignment brief belongs to the University of Melbourne and is not reproduced on this site; the task is paraphrased instead. If you are taking this subject, please do your own work.