← Gym/Confusion Matrix Metrics
00:00/ 25 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Implement classification_metrics(y_true, y_pred) — the confusion matrix and its derived metrics for binary classification, without sklearn.

A staple of MLE interviews, and what separates answers is how you handle division by zero.

Input

Two lists of the same length. Each element is 0 or 1.

python
y_true = [1, 1, 0, 1, 0]
y_pred = [1, 0, 0, 1, 1]

Confusion matrix

The positive class is 1.

text
TP = actual 1, predicted 1        FN = actual 1, predicted 0
FP = actual 0, predicted 1        TN = actual 0, predicted 0

Return

python
{"tp": 2, "tn": 1, "fp": 1, "fn": 1,
 "precision": 0.666667, "recall": 0.666667, "f1": 0.666667}

Metric definitions

text
precision = TP / (TP + FP)          ← denominator is "what we predicted positive"
recall    = TP / (TP + FN)          ← denominator is "what actually is positive"
f1        = 2·precision·recall / (precision + recall)

Non-negotiables

  • Never divide by zero. When a denominator is zero, that metric is 0.0.
    • Nothing predicted positive → precision = 0.0
    • No actual positives → recall = 0.0
    • precision + recall == 0 → f1 = 0.0
  • Round all three metrics to 6 decimal places (round(x, 6)).
  • Empty input ([], []) gives four zero counts and three 0.0 metrics.

Level 1 · Confusion matrix and metrics

Implement classification_metrics(y_true, y_pred).

  • The two lists always have the same length.
  • The returned dictionary has exactly these seven keys: tp, tn, fp, fn, precision, recall, f1.
  • Counts (tp and friends) are ints; metrics are floats.