🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Implement a bagging classifier over decision stumps, without sklearn. Half of this problem is reproducibility — the same seed must always produce the same result, so the random-number protocol has to be followed exactly.
A depth-1 tree. Its best split is found by minimizing weighted Gini.
X[:, f] <= t, right = the restmodel = BaggingClassifier(n_estimators=10, seed=42)
model.fit(X, y)
pred = model.predict(X) # a list of n 0/1 values
rng = np.random.default_rng(seed)
for i in range(n_estimators):
idx = rng.integers(0, n_samples, size=n_samples) # sample with replacement
# fit one stump on X[idx], y[idx]
Create rng once and make the call above exactly once per estimator, in
order. Deviate from this and the same seed yields different results.
Collect the stumps' predictions and take a majority vote. 0 on a draw.
y is 0 or 1.
Build the base learner first.
fit_stump(X, y) -> dict
It returns one of these two shapes.
{"feature": int, "threshold": float, "left": 0|1, "right": 0|1}
{"feature": None, "threshold": None, "left": c, "right": c} # no split possible
feature and threshold are None, and both left
and right hold the overall majority class (0 on a draw).predict_stump(stump, X) -> list
left when X[:, feature] <= threshold, otherwise right.feature is None, return the left value for every row.