← Gym/Logistic Regression from Scratch
00:00/ 25 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Train a binary classifier that is linear regression with one sigmoid bolted on, without sklearn. The skeleton matches [[algo/mle-interview/linear-regression-scratch|linear regression]] — only the prediction function and the gradient constant differ.

Input

python
X  # (n, d) real-valued features. A list of lists or a numpy array
y  # (n,)  0 or 1

If X arrives one-dimensional, treat it as (n, 1).

Model

Prepend a column of ones to the design matrix — w[0] is the intercept and the weight vector has length d + 1.

σ(z)=11+e−z,p^=σ(Aw)\sigma(z) = \frac{1}{1 + e^{-z}}, \qquad \hat{p} = \sigma(Aw)

Training rule

Exactly two things differ from linear regression. Miss either one and every number changes.

  • The prediction is sigmoid(Aw), not Aw.

  • The log-loss gradient has no constant 2:

    ∇wL=1nA⊤(σ(Aw)−y)\nabla_w L = \frac{1}{n}A^\top(\sigma(Aw) - y)

Everything else is the same — start from the zero vector and apply w = w - lr * grad exactly epochs times.

Return

  • fit_logistic_regression(X, y, lr, epochs) → a list of length d+1
  • Round every float to 6 decimal places.

There is no randomness — the same input always produces the same numbers.

Level 1 · Log-loss gradient descent

Implement fit_logistic_regression(X, y, lr, epochs).

  • Return a list of length d+1, w[0] is the intercept, start from the zero vector.
  • The prediction is sigmoid(Aw) and the gradient is A.T @ (sigmoid(Aw) - y) / n — no constant 2. Copying it straight from linear regression is exactly where this goes wrong.
  • Round each element with round(v, 6).