🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.
A pipeline that turns ad logs into features for training a CTR model. This is the problem closest to an ads ML engineer's actual work, and it confronts data leakage head-on.
rows = [
{"user": "u1", "ad": "a1", "device": "ios", "click": 1},
{"user": "u1", "ad": "a2", "device": "ios", "click": 0},
{"user": "u2", "ad": "a1", "device": "android", "click": 1},
]
Each row is one impression. click is 0 or 1.
Three features for each of the three axes (user, ad, device) — nine in total,
appended to the original row.
{axis}_impressions impressions for that axis value
{axis}_clicks clicks for that axis value
{axis}_ctr clicks / impressions (0.0 when the denominator is 0)
user, ad, device, click) and add the nine.round(x, 6)).| Level | Statistics scope |
|---|---|
| 1 | aggregated over the whole dataset |
| 2 | only rows before this one (leakage-free) |
Level 2 is the heart of this problem.
Implement build_ctr_features(rows).
Each row's statistics count every impression where that axis value appears anywhere in the dataset.
rows = [
{"user": "u1", "ad": "a1", "device": "ios", "click": 1},
{"user": "u1", "ad": "a2", "device": "ios", "click": 0},
{"user": "u2", "ad": "a1", "device": "android", "click": 1},
]
build_ctr_features(rows)[0]
# {"user": "u1", "ad": "a1", "device": "ios", "click": 1,
# "user_impressions": 2, "user_clicks": 1, "user_ctr": 0.5,
# "ad_impressions": 2, "ad_clicks": 2, "ad_ctr": 1.0,
# "device_impressions": 2, "device_clicks": 1, "device_ctr": 0.5}
Two passes are enough — one to aggregate, one to attach.