🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Add a penalty to linear regression to keep the weights small. The step people get wrong is not the penalty term itself — it is leaving the intercept out of it.
The intercept shifts every prediction up or down as a block, so it says nothing
about model complexity. Penalize it and the model becomes systematically biased
the further the target mean sits from zero. That is why standard implementations
always exclude w[0] from the penalty.
Same as [[algo/mle-interview/linear-regression-scratch|linear regression]] —
prepend a column of ones to the design matrix, start from the zero vector, and
apply w = w - lr * grad exactly epochs times.
The base gradient is unchanged:
w' is the weight vector with the intercept slot zeroed out.
The absolute value is not differentiable at zero, so use a subgradient.
Lasso's coefficient carries no factor of 2 — that is the difference from ridge.
fit_ridge(X, y, lr, epochs, alpha) / fit_lasso(X, y, lr, epochs, alpha)d+1, every float rounded to 6 decimal places.alpha=0 both must reproduce plain linear regression exactly.Implement fit_ridge(X, y, lr, epochs, alpha).
epochs times.grad = (2/n) A.T @ (A @ w - y) + 2 * alpha * w'w' is the weight vector with the intercept slot zeroed — no penalty lands
on w[0].round(v, 6).Get alpha=0 matching plain linear regression first; then you can verify the
penalty term on its own.