Skip to content
Source Separation Lab

Methods · how every number on this site is made

Methods, decisions and limits

Where the data come from, which model generates them, what each estimator computes, how the 2026 analyses put uncertainty on the results, and where all of it stops being trustworthy.

01

Data provenance

Every dataset on this site is synthetic. There are no people, no personal information and no real measurements. Six time courses and six spatial maps are generated from fixed rules, mixed, and buried in Gaussian noise drawn from a seeded random generator.

The generator is a TypeScript port of numpy's legacy RandomState, so seed 30034 reproduces the noise of the 2021 notebook bit for bit (decision record DR-002). The parity tests compare the port against the files the 2021 notebook wrote to coursework/stage/. The 2026 statistics are checked against numpy, scipy and scikit-learn through scripts/export_stats_reference.py.

The repository also keeps a set of noise drawn in R with the same seed. The final notebook did not use it. A test recomputes the 2021 pipeline on it, and the methods rank in the same order with numbers that differ in the second or third decimal place.

02

Model

The data are a product of temporal and spatial sources, each with its own noise, with N = 240 time points and V = 441 pixels:

X=(TC+Γt)(SM+Γs),Γt∈R240×6,Γs∈R6×441.\begin{gathered} X = (\mathrm{TC} + \Gamma_t)(\mathrm{SM} + \Gamma_s), \\ \Gamma_t \in \mathbb{R}^{240\times 6},\quad \Gamma_s \in \mathbb{R}^{6 \times 441}. \end{gathered}

Expanding the product shows which terms act as signal and which as error:

X=TC SM⏟DA+TC Γs+Γt SM+Γt Γs⏟E.X = \underbrace{\mathrm{TC}\,\mathrm{SM}}_{DA} + \underbrace{\mathrm{TC}\,\Gamma_s + \Gamma_t\,\mathrm{SM} + \Gamma_t\,\Gamma_s}_{E}.

The regression model is X=DA+EX = DA + E with D=TCD = \mathrm{TC} known. Each column of X is then standardised, so the estimand is the map on the standardised scale rather than the 0/1 map that generated it.

03

Assumptions

  • Gaussian, independent noise. Every entry of Γt and Γs is drawn independently from a normal distribution with mean 0 and variance 0.25 (temporal) or 0.015 (spatial), and the two noise matrices are independent of each other and of the sources.
  • Known regressors. The estimators are given the true time courses. Real source separation has to estimate them too.
  • Standardisation choices. Time courses are standardised with the sample standard deviation (ddof = 1, as pandas does). Columns of X are standardised with the population standard deviation (ddof = 0, as scikit-learn's StandardScaler does). The spatial maps are not standardised. A constant column of X is centred and left unscaled.
  • Non-overlapping patches. Each pixel belongs to at most one source, which is why the noise-free target is the true map scaled by N/(N−1)\sqrt{N/(N-1)}.
  • Penalties as submitted. Ridge applies λ to D⊤DD^\top D without the brief's scaling by V, and the 2021 lasso is a single thresholded step. Both quirks are kept so the report's numbers reproduce.

04

Estimators

All five estimators return maps  and time courses D̂ = XÂᵀ.

Least squares and ridge have closed forms:

A^LSR=(D⊤D)−1D⊤X,A^RR=(D⊤D+λI)−1D⊤X.\begin{gathered} \hat A_{\mathrm{LSR}} = (D^\top D)^{-1} D^\top X, \\ \hat A_{\mathrm{RR}} = (D^\top D + \lambda I)^{-1} D^\top X . \end{gathered}

The lasso as submitted takes one soft-thresholded gradient step from zero with step size s=1/(1.1 ∥D⊤D∥F)s = 1/(1.1\,\lVert D^\top D\rVert_F) and threshold t=Nρst = N\rho s, then scales the survivors by 1 + t (the (1 / 1 + thr) quirk):

A^LR=(1+t) sign⁡(sD⊤X)×max⁡ ⁣(∣sD⊤X∣−t, 0).\begin{aligned} \hat A_{\mathrm{LR}} = {} & (1+t)\,\operatorname{sign}(s D^\top X) \\ & \times \max\!\big(\lvert s D^\top X\rvert - t,\ 0\big). \end{aligned}

The converged lasso (added in 2026, decision record DR-003) solves the objective the 2021 code was approximating, pixel by pixel, by cyclic coordinate descent with G=D⊤DG = D^\top D and c=D⊤xvc = D^\top x_v:

a^v=arg⁡min⁡a12N∥xv−Da∥22+ρ∥a∥1,aj←S(cj−∑k≠jGjkak, Nρ)Gjj.\begin{gathered} \hat a_v = \arg\min_a \tfrac{1}{2N}\lVert x_v - D a\rVert_2^2 + \rho\lVert a\rVert_1, \\ a_j \leftarrow \frac{S\big(c_j - \sum_{k\ne j} G_{jk} a_k,\ N\rho\big)}{G_{jj}}. \end{gathered}

Principal component regression runs the 2021 lasso on the left singular vectors Z=UZ = U of TC=UΣW⊤\mathrm{TC} = U\Sigma W^\top with ρ = 0.001. The singular values of TC run from 24.20 down to 6.05.

05

Evaluation design

Metrics. Reconstruction error is ∥X−D^A^∥F2/NV\lVert X - \hat D \hat A\rVert_F^2 / NV, the criterion Question 2.3 used to choose ρ. Recovery is the absolute correlation between a true source and a retrieved one. The 2021 sums ΣcT and ΣcS pair sources by index. The 2026 per-source scores take the maximum over retrieved sources, as the brief defines them, and report how often that maximum is the matching index. Bias² and variance are measured against the noise-free target A∗=(D⊤D)−1D⊤X0A^\ast = (D^\top D)^{-1}D^\top X_0 and divided by ∥A∗∥2\lVert A^\ast\rVert^2.

Realisations and seeds. Every Monte-Carlo draws from numpy's generator with the seed shown on the page (30034 by default). In the 2026 analyses, realisation 1 is the 2021 dataset itself. On /tune the as-submitted design continues the 2021 stream, so realisations 1 to 10 are the ten behind the report.

Paired comparisons. Whenever two estimators or two penalties are compared, they are fitted to the same datasets (common random numbers). The site reports the mean of the per-realisation differences with a 95% interval, the standardised effect size dz=Δˉ/sΔd_z = \bar\Delta / s_\Delta, how often one beats the other with a Wilson interval, and an exact two-sided sign test. The p-value is never reported alone. Because paired differences can be almost constant, dzd_z measures how consistent a difference is, not how large it is, so the size of each difference is also given relative to the comparison method's mean, and differences under 1% are called negligible. A comparator is fixed in advance wherever possible. On /tune the cost of ρ = 0.625 is measured against ρ = 0.60, the lowest point of the 2021 curve, rather than against the best ρ of the same run, which is picked from the same datasets and so makes the gap look larger than it is (DR-004).

Intervals. From 30 realisations up, means use 95% percentile-bootstrap intervals over realisations with 2,000 resamples and resampling seed 30034. The resampling indices come from the same numpy generator, so each interval can be reproduced exactly in Python. Below 30 realisations the percentile bootstrap of a mean runs narrow: in a simulation with normal data it covered the true mean 90% of the time at n = 10 and 92% at n = 20. There, means use the Student-t interval xˉ±tn−1, 0.975 s/n\bar x \pm t_{n-1,\,0.975}\, s/\sqrt n instead, and captions say so (DR-005). Where the best ρ lands is always a bootstrap, because it is not a mean. Proportions use the Wilson score interval. Every helper lives in web/src/lib/stats/ and is tested against numpy and scipy values.

Sensitivity. The noise grid restarts the seeded stream in every cell, so all cells share the same standard-normal draws scaled to their own variances. Differences between cells then come from the noise level, not from a different random draw.

06

Limitations

  • Synthetic setting. Boxcar time courses, square patches and Gaussian noise are much cleaner than real fMRI. Nothing here says how these estimators behave on real recordings.
  • Multicollinearity by design. TC4, TC5 and TC6 overlap in time, with correlations up to 0.77. The condition number of TC is 4.0, which inflates the variance of least squares and is what the penalties are meant to fix.
  • Known time courses. Every estimator is given D = TC. In practice the time courses are unknown, and methods such as independent component analysis would be needed.
  • What the intervals mean. The intervals describe Monte-Carlo uncertainty for this simulation design. They are not confidence intervals for any single real dataset, and they shrink as you add realisations. On step 05 only the MSE carries an interval; its two parts, bias² and variance, are shown as point estimates.
  • The tuning criterion. The 2021 MSE reconstructs X from itself (D̂ = XÂᵀ), so it rewards fitting X and only indirectly measures recovery of the sources.
  • Plug-in bias². The plug-in bias² is too high by about the variance divided by R − 1, because the average over R realisations is itself noisy. At 50 realisations that is 0.033 for least squares on the maps, two thirds of its plug-in bias² of 0.050, so step 05 also shows a debiased bias² (0.018 here). For ridge and both lassos the correction is 0.005 or less.
  • Two scoring conventions. In 2021 the spatial maps were scored against a copy in a different pixel order from the one X was built with. Step 02 keeps that, so it reproduces the report. The 2026 analyses use the matching order, so their spatial sums are higher.

07

What I'd change

  • Choose ρ with a paired design, more realisations and an interval on the minimiser, and refine the grid near the minimum instead of taking a midpoint (DR-001).
  • Tune penalties for the quantity of interest, recovery of the true sources, rather than a reconstruction of X. The converged lasso's map error is lowest near ρ = 0.40, not 0.625.
  • Use a converged lasso from the start and test it against a reference solver (DR-003).
  • Apply ridge with the brief's scaling, λ̃ = λV. The bias-variance view puts the optimum near λ ≈ 300.
  • Let the penalty scale with the noise level, so it still works when the noise changes.
  • Write comparators and the interval method into an analysis plan before running a simulation (DR-004, DR-005).
  • Save the exact noise next to the outputs, and generate data in one language only (DR-002).

08

Decision records

Each record states the decision first, then the context, the options, the reasons, what happened with the real numbers, weak ones included, and what I would change. Past records are never edited. A new record supersedes an old one.

  1. DR-001 · 2026-10-06

    Choosing the lasso penalty ρ by the mean MSE over realisations

    I chose ρ = 0.625, the midpoint of 0.60 and 0.65, from the mean reconstruction MSE over ten noise realisations, and the revival keeps that value wherever it reproduces the report.

  2. DR-002 · 2026-10-06

    Which noise to reproduce for parity: numpy's stream, with the R-generated noise kept as a check

    The browser port regenerates the noise the 2021 Python notebook actually used, by porting numpy's legacy random generator bit for bit, and keeps the R-generated noise in coursework/data/ as provenance and as a parity check, not as the source of the reported numbers.

  3. DR-003 · 2026-10-06

    Adding a converged lasso by coordinate descent next to the as-submitted one

    The 2026 analyses add a lasso solved to convergence by cyclic coordinate descent, verified against scikit-learn, while the as-submitted single-step lasso stays wherever the report's numbers are reproduced.

  4. DR-004 · 2026-10-06

    Measuring the cost of ρ = 0.625 against a comparator fixed in advance

    On /tune the cost of ρ = 0.625 is measured against ρ = 0.60, the lowest point of the 2021 mean curve, which was fixed before any 2026 data were drawn; the comparison with the best ρ of the current run is kept only as a secondary figure, labelled as biased upwards.

  5. DR-005 · 2026-10-06

    Student-t intervals for means over fewer than 30 realisations

    Means over 30 or more Monte-Carlo realisations keep the 95% percentile-bootstrap interval, and means over fewer than 30 use the Student-t interval, because the percentile bootstrap undercovers at small n.

09

Model card

This card covers the five regression estimators the site uses to retrieve the hidden sources: least squares (LSR), ridge (RR), the lasso as submitted in 2021 (LR), the converged lasso added in 2026 (LR-CD) and principal component regression (PCR). They are statistical estimators on a synthetic teaching dataset, not trained models deployed on real data. The card follows the usual model-card headings so their limits are written down in one place.

Model details

  • Author: Sunchuangyu (Rin) Huang. The estimators were written for MAST30034 Applied Data Science, University of Melbourne, Semester 2 2021, and ported to TypeScript in 2026.
  • Task: given data X (240 time points × 441 pixels) and known time courses D = TC (240 × 6), estimate the spatial maps A (6 × 441) and then the time courses D̂ = X Âᵀ.
  • Estimators:
    • LSR: A = (DᵀD)⁻¹DᵀX.
    • RR: A = (DᵀD + λI)⁻¹DᵀX, with λ applied unscaled as in 2021 (λ = 0.5).
    • LR (as submitted): one soft-thresholded gradient step with step 1/(1.1 ∥D⊤D∥F)1 / (1.1\,\lVert D^\top D\rVert_F), threshold N ρ × step and the (1 / 1 + thr) scaling quirk (ρ = 0.625).
    • LR-CD (2026): the same lasso objective, 12N∥xv−Da∥22+ρ∥a∥1\tfrac{1}{2N}\lVert x_v - D a\rVert_2^2 + \rho \lVert a\rVert_1, solved to convergence by coordinate descent (ρ = 0.625 unless stated).
    • PCR: LR applied to the left singular vectors of TC (ρ = 0.001).
  • Implementation: framework-free TypeScript in web/src/lib/regression.ts, run in the browser.

Intended use

  • Teaching and portfolio use. The site shows how penalties trade bias against variance, how to put intervals on Monte-Carlo results, and how to compare estimators fairly.
  • Reproducing the 2021 report exactly, and extending it with uncertainty without changing its numbers.

Out of scope

  • Any real fMRI or neuroimaging data. The generative model is a toy, and nothing here has been validated on real recordings.
  • Settings where the time courses are unknown. All five estimators assume D is given, which real source separation does not.
  • Any clinical, diagnostic or decision-making use.

Data provenance

  • All data are synthetic and contain no personal information. Six boxcar time courses and six square spatial maps on a 21 × 21 grid are generated from fixed timings, then mixed as X = (TC + Γt)(SM + Γs) with Gaussian noise of variance 0.25 (temporal) and 0.015 (spatial), and each column of X is standardised.
  • The noise is regenerated from the seed by a port of numpy's legacy RandomState. With seed 30034 the browser reproduces the 2021 noise bit for bit (DR-002).
  • The parity oracle is the set of CSVs the 2021 notebook wrote to coursework/stage/. Reference values for the 2026 statistics come from numpy, scipy and scikit-learn via scripts/export_stats_reference.py.

Evaluation

Evaluation uses repeated noise realisations from seed 30034. Realisation 1 is the 2021 dataset. Intervals for means are 95% percentile bootstraps over realisations (2,000 resamples, seed 30034) from 30 realisations up, and Student-t intervals below 30, where the percentile bootstrap runs narrow (DR-005). Proportions use Wilson intervals. Spatial maps are scored against the layout X was built with (the aligned convention). The as-submitted scoring is kept on /retrieve, where it reproduces the report.

Recovery over 100 realisations, as Σ over six sources of |corr| with the matched retrieved source (maximum 6):

Estimator ΣcT\Sigma c_T (time courses) ΣcS\Sigma c_S (maps, aligned)
LSR 5.217 (5.207 to 5.226) 3.312 (3.305 to 3.318)
RR, λ = 0.5 5.221 (5.212 to 5.231) 3.313 (3.306 to 3.319)
LR as submitted, ρ = 0.625 5.435 (5.432 to 5.439) 4.825 (4.813 to 4.837)
LR-CD, ρ = 0.625 5.419 (5.415 to 5.423) 5.146 (5.136 to 5.157)
PCR, ρ = 0.001 0.664 (0.637 to 0.690) 0.480 (0.468 to 0.491)

Paired comparisons on the same 100 realisations confirm the 2021 claims. LR beats RR on ΣcT\Sigma c_T by 0.214 (0.205 to 0.223) and on ΣcS\Sigma c_S by 1.512 (1.498 to 1.527), 46% of ridge's score, in 100 of 100 realisations each. RR beats LSR on ΣcT\Sigma c_T in 100 of 100 realisations, but only by 0.005 (0.0048 to 0.0050), about 0.1% of the least-squares score, which is negligible. Because both estimators see the same noise, such a difference can be very consistent (dzd_z = 7.0) while being far too small to matter, so the site reports the relative size next to dzd_z.

Error against the noise-free target over 50 realisations, relative to the size of the target (1.0 is the error of predicting all zeros):

Estimator Best penalty Maps: MSE (bias², variance) Time courses: best penalty and MSE
LSR none 1.647 (0.050, 1.596) 2.950
RR λ ≈ 316 0.725 (0.483, 0.242) λ ≈ 562, 0.205
LR-CD ρ = 0.40 0.437 (0.293, 0.144) ρ = 0.40, 0.190
LR as submitted ρ = 0.25 0.726 (0.662, 0.064) ρ = 0.15, 0.267

The MSE has a 95% bootstrap interval on the site, for example 0.725 (0.724 to 0.726) for ridge at its best λ and 1.647 (1.633 to 1.660) for least squares. Bias² and variance are point estimates of its two parts, without intervals. The bias² shown is the plug-in estimate, which adds up with the variance to the MSE but is inflated by Monte-Carlo noise of about variance / (R − 1). For least squares on the maps that inflation is 0.033 of its 0.050, which leaves a debiased bias² of 0.018. For ridge and both lassos it is 0.005 or less. The site shows both values.

These numbers are for the default settings and seed. The site recomputes them live, so they change if you change the dataset on step 1.

Known failure modes

  • Multicollinearity by design. TC4, TC5 and TC6 are correlated (|r| up to 0.77), and the condition number of TC is 4.0. Least squares is nearly unbiased for the maps (debiased bias² 0.018; the plug-in value is 0.050) but its variance is large (1.60), about 90 times its bias², which is what the penalties are for.
  • Degenerate timings. If two time courses coincide, DᵀD is singular. LSR is then undefined and the site reports this instead of failing.
  • A fixed penalty does not transfer across noise levels. With ρ = 0.625 and temporal noise variance raised to 2, both lassos zero almost every coefficient and the mean recovery of the maps drops to about 0.1. Ridge and LSR degrade smoothly instead.
  • The as-submitted lasso is not converged. It is heavily biased at every ρ (bias² above 0.59 for the maps) and should not be used to reason about the lasso in general (DR-003).
  • PCR loses the sources. Principal components mix the time courses, so no retrieved component lines up with a single source, and PCR scores lowest on the matched-index sums ΣcT\Sigma c_T and ΣcS\Sigma c_S (0.66 and 0.48 out of 6). It is not lowest on every measure: on the maximum-correlation score for TC5 it reaches 0.892 (0.890 to 0.894), above least squares (0.838) and ridge (0.841).
  • The ρ selection criterion reconstructs X from itself. It rewards fitting X and is only indirectly related to recovering the sources (DR-001).

Ethical considerations

  • No people, personal data or real measurements are involved. The main risk is over-generalisation, reading results from a clean toy model as statements about real brain data. The methods page and this card state the synthetic setting, and the instructions given to the optional AI feature rule out real-world or clinical conclusions.
  • Academic integrity. The original coursework is published as a record. The site paraphrases the brief instead of reproducing it and asks current students not to reuse the work.
  • The optional AI feature only explains figures from numeric summaries. It never produces or changes any number in this card. See the AI use statement on /methods.

Caveats and recommendations

  • Intervals describe Monte-Carlo uncertainty for this simulation design. They are not confidence intervals for any real dataset.
  • The plug-in bias² is biased upwards by about the variance divided by R − 1. At R = 50 that is 0.033 for LSR on the maps, two thirds of its plug-in bias², so the site also reports a debiased bias². For the other estimators in the table the correction is 0.005 or less on the maps.
  • If you adapt these estimators, tune the penalty for the target you care about and for your own noise level, and use a converged solver.

10

AI use statement

What AI does here. One optional feature, “Explain this figure”, on steps 03, 05 and 06. It sends a short numeric summary of the figure on screen to a model and shows back a plain-English reading with caveats. The site is complete without it.

What it never does. It never computes, changes or replaces a number, a figure or a conclusion on the site. Every statistic comes from the deterministic code described above. Its prompt never contains the underlying data matrices, other pages or any personal data. It never runs without a click, and it never uses a key belonging to this site, because there is none.

Your key. You bring your own Anthropic or OpenAI key. It is kept in this browser's sessionStorage, or in localStorage only if you tick “remember on this device”, and a “forget key” button removes it. The browser sends it straight to the provider you chose. It is never sent to this site's server, never logged and never written to the audit log. The default Anthropic model is Claude Haiku 4.5 (claude-haiku-4-5), with Claude Sonnet 5.5 as an option. For OpenAI the model id is editable (default gpt-5-mini).

Data sent to the provider. Fixed instructions, plus a JSON summary of at most a few kilobytes holding the figure's title, a fixed description, and rounded numbers such as means, intervals, sample sizes and seeds. You can read the exact text that would be sent, instructions and summary, under each button before sending it. Because the request goes straight from your browser, the provider also receives the usual request metadata: your IP address, your browser's user agent and this site's origin. The call is billed to and logged against your own provider account, and the provider's terms govern what happens to it there.

Checks and human review. Answers must match a fixed JSON schema, which is validated in the browser. An automatic check compares every number in the answer with the numbers in the summary and flags any that do not match. Counts such as “7 of 50” and percentages must match a count or share in the summary exactly. The check shows that a number appears in the summary, not that it is attached to the right quantity, and it does not check years or the 95% confidence level. Every answer is labelled “AI-generated” and asks you to accept, edit or reject it.

Audit trail. Each call is recorded in this browser's IndexedDB with the exact request sent (endpoint, API version, model, token limit, output schema, instructions and summary, but never the key), the output, latency, token usage, check results and your decision. Failed and cancelled calls are recorded too, and when the model replies in a form that fails validation its raw reply is kept. An explanation is shown only next to the figure it was generated for; if the figure changes, it is hidden until the figure matches again. You can read and export it as JSON or CSV at /ai-log, and it never leaves your device unless you export it.

Frameworks. The design is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It does not claim compliance with, or certification under, any of them.