← Gym/MLP Backprop from Scratch
00:00/ 35 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Write the forward and backward pass of a one-hidden-layer network by hand. The weights are handed to you, so there is no random initialization — the only thing under test is whether you differentiated correctly.

Architecture

For a single sample x of length d:

z1=W1x+b1,a1=σ(z1),y^=W2⋅a1+b2z_1 = W_1 x + b_1, \quad a_1 = \sigma(z_1), \quad \hat{y} = W_2 \cdot a_1 + b_2

  • σ\sigma is the sigmoid, applied to the hidden layer only. The output layer is linear.
  • The loss is squared error L=(y^−y)2L = (\hat{y} - y)^2 — one sample, so there is no mean.

Parameters

python
params = {
    "W1": [[...], ...],   # (h, d)
    "b1": [...],          # (h,)
    "W2": [...],          # (h,)
    "b2": 0.0,            # scalar
}

Backward pass

Follow the chain rule. The key fact is that the sigmoid's derivative is σ(1−σ)\sigma(1-\sigma).

∂L∂y^=2(y^−y)\frac{\partial L}{\partial \hat{y}} = 2(\hat{y} - y)

\frac{\partial L}{\partial b_2} = \frac{\partial L}{\partial \hat{y}}$$ $$\frac{\partial L}{\partial z_1} = \frac{\partial L}{\partial \hat{y}} W_2 \odot a_1(1 - a_1)$$ $$\frac{\partial L}{\partial W_1} = \frac{\partial L}{\partial z_1} x^\top, \qquad \frac{\partial L}{\partial b_1} = \frac{\partial L}{\partial z_1}$$ **Do not drop the constant 2** — every gradient comes out exactly half if you do. ## Return Round every float to **6 decimal places**.

Level 1 · Forward pass

Implement forward(params, x).

  • z1 = W1 @ x + b1, a1 = sigmoid(z1), output = W2 @ a1 + b2
  • The sigmoid goes on the hidden layer only — the output layer is linear. Squashing the output too throws off every value downstream.
  • Return {"z1": [...], "a1": [...], "output": scalar}, every float round(v, 6).