🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
Build a classifier that does no training at all. The hard part is not the distance calculation — it is how you break ties, which is where implementations diverge.
X_train # (n, d) real-valued features
y_train # (n,) integer labels
X_query # (m, d) points to predict
k # neighbor count (at least 1)
If X_train or X_query arrives one-dimensional, treat it as (n, 1) / (m, 1).
Use Euclidean distance, square root included.
Ties decide the answer, so follow these exactly.
(distance, training index) ascending and take the first k.
On equal distance the lower index comes first.k labels. On a tie, pick the smaller label value.k exceeds the number of training samples, use all of them.knn_predict(X_train, y_train, X_query, k) → a list of m labels.
Implement knn_predict(X_train, y_train, X_query, k).
k by (distance, index) ascending.k exceeds the training set size, use all of it.m labels.