🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Build a classifier out of a single (mostly false) assumption: that the features are independent. What makes this model distinctive is that training is one pass of aggregation — no iterations, no learning rate.
X # (n, d) real-valued features
y # (n,) integer labels
If X arrives one-dimensional, treat it as (n, 1).
Handle classes in ascending sorted order. For class c:
ddof=0, i.e. divide by n_c) A zero variance would divide by zero during prediction, so add 1e-6 to every
variance.
The usual choice is 1e-9, and there is a reason it is 1e-6 here. Rounding to
6 decimal places — as the return rule below requires — turns 1e-9 straight into
0.0, the smoothing disappears, and prediction blows up on log(0). The constant
has to survive the rounding.
Pick the class with the largest log posterior.
On a tie, pick the smaller class value.
fit_gaussian_nb(X, y) → a dictionary with these four keys. Round every float to
6 decimal places.
{"classes": [...], "priors": [...], "means": [[...], ...], "vars": [[...], ...]}
Implement fit_gaussian_nb(X, y).
classes is the class list in ascending sorted order.priors[j] = n_c / nmeans[j][i] and vars[j][i] — per-class, per-feature mean and
population variance (ddof=0). numpy's .var() already defaults to
ddof=0, so it works as-is.1e-6 to every variance — with 1e-9 the smoothing vanishes under
round(v, 6).round(v, 6).