← Gym/Run-Length Encoder
00:00/ 12 min

🧭 Do not search for the first 15 minutes. When stuck: re-read the requirements → define I/O → choose the data structure → trace a small example by hand → write code.

Run-length encoding shortens a stretch of the same repeated character into "the character plus a count". Implement compress(s), which compresses a string this way.

Input

A single arbitrary string.

python
s = "aaabbccccdaa"

Digits and spaces may appear alongside letters. None of them get special treatment — each is just a character.

Output

Return the compressed string.

python
compress("aaabbccccdaa")
# "a3b2c4d1a2"

Splitting the input into runs from the left:

  • aaa: a repeated three times → a3
  • bb: b repeated twice → b2
  • cccc: c repeated four times → c4
  • d: d appeared once → d1
  • aa: a repeated twice again → a2

Concatenated, that is "a3b2c4d1a2".

Where a run ends

A run is a stretch of the same character in a row. Separated occurrences of the same character are not merged.

In the example a appears five times overall, but the answer is not a5. The leading aaa and the trailing aa have other characters between them, so they are different runs. That is why a appears twice in the result.

Things to watch

  1. A run of length 1 still gets a number. A lone d is d1, not d.
  2. A count of 10 or more becomes a multi-digit number. Twelve a's in a row give a12.
  3. Return the result even when it is longer than the original. Deciding whether compression was worthwhile is not this function's job. "abc" becomes "a1b1c1".
  4. The final run is easy to lose. If you only append to the result when the character changes, the last run is still pending when the input ends. Getting "a2" for "aab" is that bug.
  5. An empty string returns an empty string.

Level 1 · Compression

Implement compress(s). The input is an arbitrary string (digits and spaces included).