🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Train a linear regression with MSE gradient descent, without sklearn. Every gradient-descent model is built on the skeleton you write here: initialize → predict → gradient → update, repeat.
X # (n, d) real-valued features. A list of lists or a numpy array
y # (n,) real-valued targets
If X arrives one-dimensional, treat it as (n, 1).
To learn an intercept, prepend a column of ones to the design matrix.
X = [[1], [2]] → A = [[1, 1],
[1, 2]]
So the weight vector w has length d + 1, and w[0] is the intercept.
These three details determine the answer. Get any one wrong and the numbers drift.
Start from the zero vector — w = [0, 0, ..., 0]
The loss is MSE and its gradient carries the constant 2:
Apply w = w - lr * grad exactly epochs times.
fit_linear_regression(X, y, lr, epochs) → a list of length d+1There is no randomness anywhere in this problem — the data has no noise and the initial weights are fixed, so the same input always produces the same numbers.
Implement fit_linear_regression(X, y, lr, epochs).
d+1 where w[0] is the intercept.epochs times.lr.round(v, 6).