Skip to content
BytePatterns

Encode and Decode Strings

Strings: lesson 11 of 11

Send the length first and no character is special.

Lesson 11 of 11 · 5 min

Encode and Decode Strings

Step 1 of 10

Two parts to send — and one of them contains a #. Any separator you reserve is a character the data may not hold.

The Idea

Joining parts with a separator breaks the moment a payload contains that separator. Prefix each part with its length instead: 5#hello. The decoder reads digits up to the marker, then copies exactly that many characters without looking at them. Empty strings survive, every byte is legal inside a payload, and decoding stays linear because nothing is ever scanned twice.

Real-World Example

Network protocols frame messages the same way. HTTP sends Content-Length before the body, so the parser knows where one message stops and the next begins without hunting for a delimiter that the body might contain.

The Code

def encode(parts):
    return "".join(f"{len(p)}#{p}" for p in parts)

def decode(s):
    out, i = [], 0
    while i < len(s):
        j = s.index("#", i)              # the length ends at the first #
        n = int(s[i:j])
        out.append(s[j + 1 : j + 1 + n])
        i = j + 1 + n                    # jump the whole payload
    return out

print(encode(["hi", "a#b", ""]))   # 2#hi3#a#b0#
print(decode("2#hi3#a#b0#"))       # ['hi', 'a#b', '']

Python

Your turn

What does this print?

print(len(encode(["ab", "c"])))

Mini quiz

1 / 3

Why is a plain comma separator unsafe here?

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.