🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Write the forward and backward pass of a one-hidden-layer network by hand. The weights are handed to you, so there is no random initialization — the only thing under test is whether you differentiated correctly.
For a single sample x of length d:
params = {
"W1": [[...], ...], # (h, d)
"b1": [...], # (h,)
"W2": [...], # (h,)
"b2": 0.0, # scalar
}
Follow the chain rule. The key fact is that the sigmoid's derivative is .
\frac{\partial L}{\partial b_2} = \frac{\partial L}{\partial \hat{y}}$$ $$\frac{\partial L}{\partial z_1} = \frac{\partial L}{\partial \hat{y}} W_2 \odot a_1(1 - a_1)$$ $$\frac{\partial L}{\partial W_1} = \frac{\partial L}{\partial z_1} x^\top, \qquad \frac{\partial L}{\partial b_1} = \frac{\partial L}{\partial z_1}$$ **Do not drop the constant 2** — every gradient comes out exactly half if you do. ## Return Round every float to **6 decimal places**.
Implement forward(params, x).
z1 = W1 @ x + b1, a1 = sigmoid(z1), output = W2 @ a1 + b2{"z1": [...], "a1": [...], "output": scalar}, every float round(v, 6).