Kankan

Guides

Characters, code points, bytes: why ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง is 1, 5 or 18

Updated ยท Lumen Lab

All of the numbers are right; they measure different things. ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง is 1 character as people see it, 5 Unicode code points, 8 UTF-16 code units and 18 bytes in UTF-8.

The pieces inside one emoji

The family emoji is three emoji glued together with an invisible joiner (ZWJ).

That is 5 code points. Written in UTF-16 they take 8 units, and in UTF-8 they take 18 bytes. On screen it fills one space, so to a reader it is 1 character.

Counted four ways

Counted with the โ€œtaken apartโ€ panel on this site.

TextCharactersCode pointsUTF-16 unitsUTF-8 bytes
a1111
รฉ (one code point)1112
รฉ (e plus a combining accent)1223
๐Ÿ˜€1124
๐Ÿ‘๐Ÿฝ (thumb plus skin tone)1248
๐Ÿ‡บ๐Ÿ‡ธ (U plus S)1248
๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง15818
๐Ÿด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ (flag of England)171428

Ad

What each number measures

  • Characters are what a reader would point at and call one letter or one emoji. Unicode calls them grapheme clusters, and the rules are in UAX #29. This is what the counters on this site report.
  • Code points are the numbered entries in Unicode. Combining accents, skin tones and joiners are code points of their own, which is how one visible character becomes several.
  • UTF-16 units are what JavaScript and several other languages count when you ask for the length of a string. A form that validates length that way sees ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘ง as 8.
  • Bytes are what gets stored or sent. UTF-8 uses 1 to 4 bytes per code point.

The same letter, two spellings

รฉ can be stored as one code point, or as a plain e followed by a combining accent. Both look identical. They differ in every column of the table except the first, and a search for one may not find the other. Software that cares usually normalises text to one form before counting. The twitter-text rules X uses do: either spelling counts as 1 there.

Which number does a limit use?

It depends on who set the limit, and the label rarely says. If an emoji uses up a limit faster than you expect, the form is counting units or bytes rather than characters. Two limits are documented well enough to work out exactly: X uses a weighted length, and SMS counts 7-bit slots or UTF-16 units.

Tap any character to see its pieces.

Open the character counter

Keep reading

Ad