Search

Why Unicode Exists (and Why Emojis Break Things)

Why Unicode Exists (and Why Emojis Break Things)

The short answer

Quick answer: Computers store only numbers, so text needs an agreed table that maps characters to numbers. For decades there were hundreds of incompatible tables, one per language or vendor, and text moved between them turned into garbage. Unicode is the single table that replaced them: it assigns one number, a code point, to every character in every writing system. UTF-8 is the usual way to store those numbers as bytes. Emoji break things because one visible symbol can be made of several code points, each taking several bytes, while a lot of code still assumes that one character equals one unit.

Also Read: Why Is Naming Things So Hard in Programming? - How To's

Before Unicode

ASCII

ASCII, standardised in the 1960s, used 7 bits for 128 characters: the English letters, digits, punctuation and some control codes. A is 65. It was enough for English and nothing else.

Code pages

A byte has 8 bits, so 128 values were spare. Everyone filled them differently: one table for Western European accents, another for Cyrillic, another for Greek, and many more. Languages with thousands of characters, such as Chinese, Japanese and Korean, needed multi-byte schemes of their own.

The consequences:

  • The same byte meant different characters depending on which table the reader assumed.
  • A file did not say which table it used.
  • You could not mix scripts in one document.

Open text with the wrong table and you get mojibake: "café" turning into "café".

What Unicode is

Unicode, first published in 1991, is one universal table. Each character gets a code point, written U+ followed by hexadecimal digits.

CharacterCode point
AU+0041
éU+00E9
ЖU+0416
中U+4E2D
😀U+1F600

The space runs from U+0000 to U+10FFFF, room for over a million code points, of which more than 150,000 are assigned. The first 128 match ASCII exactly.

Unicode says which number a character has. It does not say how to write that number as bytes. That is the job of an encoding.

UTF-8, UTF-16 and UTF-32

EncodingSize per code pointNotes
UTF-32Always 4 bytesSimple, wasteful
UTF-162 or 4 bytesUsed inside JavaScript, Java, .NET and Windows
UTF-81 to 4 bytesThe standard for files and the web

Early Unicode assumed 65,536 characters would be enough, so 16 bits per character seemed fine. It was not enough. UTF-16 handles the overflow with surrogate pairs: two 16-bit units that together represent one code point. Most emoji need them.

UTF-8, designed by Ken Thompson and Rob Pike in 1992 and specified in RFC 3629, is the one that won.

Also Read: Why Big Tech Companies Build Their Own Databases

Code point rangeBytesExample
U+0000 to U+007F1A
U+0080 to U+07FF2é
U+0800 to U+FFFF3中
U+10000 to U+10FFFF4😀

Its advantages:

  • Compatible with ASCII. Any ASCII file is already valid UTF-8.
  • Compact for text that is mostly Latin letters.
  • Self-synchronising. You can tell from any byte whether it starts a character or continues one.
  • No byte-order problem.

The large majority of web pages now use it.

Why emoji break things

"Length" has four different answers

TextWhat you seeCode pointsUTF-16 unitsUTF-8 bytes
A1111
é1112
😀1124
👍🏽1248
👨‍👩‍👧15818
"😀".length          // 2: JavaScript counts UTF-16 units
[..."😀"].length     // 1: code points
"👨‍👩‍👧".length        // 8
len("😀")                    # 1: Python counts code points
len("😀".encode("utf-8"))    # 4 bytes
len("👨‍👩‍👧")                   # 5

One symbol, many code points

What a reader sees as one character is called a grapheme cluster. Emoji are often built by combination:

  • Skin tones: a base emoji followed by a modifier. 👍 + 🏽.
  • Families and professions: several emoji joined by an invisible zero-width joiner (U+200D). 👨 + ZWJ + 👩 + ZWJ + 👧.
  • Flags: two "regional indicator" letters. 🇯 + 🇵 displays as the flag of Japan.
  • Variation selectors that choose between text and emoji presentation.

What goes wrong

  • Slicing in the middle. Truncating a string to N units can cut a surrogate pair or a joined sequence in half, leaving a broken character or the replacement symbol �.
  • Reversing a string scrambles combined characters.
  • Length limits disagree. A form allows 20 "characters"; the database column allows 20 bytes; the user typed 20 emoji.
  • Databases. MySQL's older utf8 character set stores at most 3 bytes per character, so it cannot hold emoji. The fix is utf8mb4. See SQL vs NoSQL for more on choosing a database.
  • Cursor movement and backspace that step by code unit and land inside a character.
  • Regular expressions where . matches half an emoji unless Unicode mode is on. See how regular expressions work.

It is not only emoji

Two ways to write the same thing

é can be one code point (U+00E9) or two: e followed by a combining accent (U+0301). They look identical and compare as different.

"é" === "é"                                   // false
"é" === "é".normalize("NFC")                  // true

Normalisation converts text to one consistent form (NFC or NFD). Do it before comparing, searching or storing identifiers.

Other traps

  • Case conversion depends on language. In Turkish, the capital of i is İ, not I. German ß upper-cases to SS.
  • Sorting depends on language too. Sorting by code point is not alphabetical order.
  • Lookalikes. Latin a and Cyrillic а are different code points that look the same, which enables spoofed domain names and usernames.
  • Right-to-left text such as Arabic and Hebrew, and the invisible direction controls that go with it.
  • The byte order mark at the start of some files, which can confuse parsers.

Practical rules

  1. Use UTF-8 everywhere: files, databases, HTTP, and APIs. Declare it: <meta charset=utf-8> and Content-Type: text/html; charset=utf-8.
  2. Decode bytes to text at the edge of your program, work with text inside, and encode on the way out.
  3. Never guess an encoding. Know it, or ask.
  4. Decide what "length" means for each limit: bytes for storage, grapheme clusters for what users see.
  5. Use a library to split text into user-visible characters. In JavaScript, Intl.Segmenter; elsewhere, an ICU-based library.
  6. Normalise before comparing.
  7. Use locale-aware functions for sorting and case.
  8. In MySQL, use utf8mb4.
  9. Test with awkward input: emoji with skin tones, accented names, Arabic, Chinese, and a very long combined sequence.

Like floating-point numbers and time zones, text looks simple until real-world data arrives. Language models face a related problem when they split text into tokens; see how tokenization works.

Frequently asked questions

What is the difference between Unicode and UTF-8?

Unicode is the table that assigns a number to each character. UTF-8 is one way of writing those numbers as bytes.

Why is the length of an emoji 2 in JavaScript?

JavaScript strings are sequences of UTF-16 code units, and most emoji need two.

What is mojibake?

Garbled text produced when bytes written in one encoding are read as another.

Should I use UTF-8 or UTF-16?

UTF-8 for storage and exchange. UTF-16 mostly appears as the internal string format of certain languages and platforms.

Conclusion

Unicode exists because the world's writing could not be squeezed into incompatible 256-character tables. It solved that problem thoroughly, and exposed another: a "character" is not a simple unit. Use UTF-8 throughout, be clear about whether you are counting bytes, code points or visible characters, and let a well-tested library handle the rest.

Related articles

Sources and further reading

TWT Staff

TWT Staff

Writes about Programming, tech news, discuss programming topics for web developers (and Web designers), and talks about SEO tools and techniques

Your experience on this site will be improved by allowing cookies Cookie Policy