← Gym/Feature Aggregator
00:00/ 22 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Implement aggregate_users(events), which extracts per-user performance metrics from ad impression logs. This is the classic preprocessing step behind CTR-model features.

Input

Each row is one impression.

python
events = [
    ("user1", "ad1", 1, 0),   # (user, ad, clicked, converted)
    ("user1", "ad2", 1, 1),
    ("user2", "ad1", 0, 0),
    ("user1", "ad1", 1, 1),
]
  • click and conversion are 0 or 1.
  • You may assume no row converts without a click.

Output

Return a dictionary holding five values per user.

python
{
    "user1": {
        "impressions": 3,
        "clicks": 3,
        "conversions": 2,
        "ctr": 1.0,
        "cvr": 0.666667,
    },
    "user2": {"impressions": 1, "clicks": 0, "conversions": 0, "ctr": 0.0, "cvr": 0.0},
}

Metric definitions

text
impressions = that user's row count
clicks      = sum of click
conversions = sum of conversion
ctr         = clicks / impressions
cvr         = conversions / clicks      ← the denominator is clicks, not impressions

Non-negotiables

  • Never divide by zero. A zero denominator makes that ratio 0.0. This is what happens to cvr for a user who never clicked.
  • Round ratios to 6 decimal places (round(x, 6)).
  • User keys keep their order of first appearance.

Level 1 · Per-user aggregation

Implement aggregate_users(events). The metric definitions and rounding rules match the shared spec.

  • An empty list returns an empty dictionary.
  • A user may be shown many different ads, and the same ad more than once.