🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Train a binary classifier that is linear regression with one sigmoid bolted on, without sklearn. The skeleton matches [[algo/mle-interview/linear-regression-scratch|linear regression]] — only the prediction function and the gradient constant differ.
X # (n, d) real-valued features. A list of lists or a numpy array
y # (n,) 0 or 1
If X arrives one-dimensional, treat it as (n, 1).
Prepend a column of ones to the design matrix — w[0] is the intercept and the
weight vector has length d + 1.
Exactly two things differ from linear regression. Miss either one and every number changes.
The prediction is sigmoid(Aw), not Aw.
The log-loss gradient has no constant 2:
Everything else is the same — start from the zero vector and apply
w = w - lr * grad exactly epochs times.
fit_logistic_regression(X, y, lr, epochs) → a list of length d+1There is no randomness — the same input always produces the same numbers.
Implement fit_logistic_regression(X, y, lr, epochs).
d+1, w[0] is the intercept, start from the zero vector.sigmoid(Aw) and the gradient is A.T @ (sigmoid(Aw) - y) / n —
no constant 2. Copying it straight from linear regression is exactly where
this goes wrong.round(v, 6).