← Gym/Linear Regression from Scratch
00:00/ 25 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Train a linear regression with MSE gradient descent, without sklearn. Every gradient-descent model is built on the skeleton you write here: initialize → predict → gradient → update, repeat.

Input

python
X  # (n, d) real-valued features. A list of lists or a numpy array
y  # (n,)  real-valued targets

If X arrives one-dimensional, treat it as (n, 1).

Model

To learn an intercept, prepend a column of ones to the design matrix.

text
X = [[1], [2]]  →  A = [[1, 1],
                        [1, 2]]

So the weight vector w has length d + 1, and w[0] is the intercept.

Training rule

These three details determine the answer. Get any one wrong and the numbers drift.

  • Start from the zero vector — w = [0, 0, ..., 0]

  • The loss is MSE and its gradient carries the constant 2:

    L(w)=1n∥Aw−y∥2,∇wL=2nA⊤(Aw−y)L(w) = \frac{1}{n}\lVert Aw - y \rVert^2, \qquad \nabla_w L = \frac{2}{n}A^\top(Aw - y)

  • Apply w = w - lr * grad exactly epochs times.

Return

  • fit_linear_regression(X, y, lr, epochs) → a list of length d+1
  • Round every float to 6 decimal places.

There is no randomness anywhere in this problem — the data has no noise and the initial weights are fixed, so the same input always produces the same numbers.

Level 1 · Gradient descent

Implement fit_linear_regression(X, y, lr, epochs).

  • Return a list of length d+1 where w[0] is the intercept.
  • Initial weights are the zero vector; update exactly epochs times.
  • Keep the constant 2 in the gradient — dropping it halves your convergence speed at the same lr.
  • Round each element with round(v, 6).
  • numpy is fine, and so is plain Python.