Step 01Question 1 · synthetic data
Building a brain scan with known answers
Real fMRI data never tells you which signals went in. So the assignment starts by inventing them: six time courses, six spatial maps, a little noise, and a mixing rule. Every later step is graded against these ground truths.
The model is the linear one at the heart of the course. Each of the V = 441 pixels records a time series of N = 240 samples, and the whole recording is a product of temporal sources D and spatial sources A plus an error term:
To generate data that follows it, noise is added to both kinds of source before they are multiplied together:
Change anything below and every figure on this page, and on the next three, updates. The defaults are the values from the assignment, and the noise is regenerated from the same seed, so out of the box you are looking at the exact matrices from my 2021 notebook.
Using the 2021 parameters. Results match the report.
Figure 1.1
Six temporal sources, standardised
TC1AV 0 · IV 30 · dur 15
TC2AV 20 · IV 45 · dur 20
TC3AV 0 · IV 60 · dur 25
TC4AV 0 · IV 40 · dur 15
TC5AV 0 · IV 40 · dur 20
TC6AV 0 · IV 40 · dur 25
Why standardise rather than normalise? In 2021 I argued that the time courses are 0/1 indicators, so dividing by the ℓ2 norm only rescales them and leaves their offset in place. Standardising centres each one (no intercept is needed) and gives it unit variance, so no source dominates the regression just because it is switched on more often.
Figure 1.2
How related are the sources to each other?
Time courses (TC)
Spatial maps (SM)
Figure 1.3
Six spatial sources on a 21 × 21 grid
SM1
SM2
SM3
SM4
SM5
SM6
sm.transpose().flatten()).Why don't the maps need standardising? They are binary masks of equal height. Rescaling them would not change which pixels belong to which source, and the time courses already carry the scale of the signal.
Figure 1.4
Gaussian noise, temporal and spatial
Temporal noise Γt ~ N(0, 0.25)
Spatial noise Γs ~ N(0, 0.015)
Correlation of Γt columns
Correlation of Γs rows
Figure 1.5
The observed dataset X = (TC + Γt)(SM + Γs)
100 sampled time series from X
Variance of each of the 441 variables
Next · Step 02
Retrieve
Recover the sources with least squares, ridge, lasso and PCR.