Characters, code points, bytes: why ๐จโ๐ฉโ๐ง is 1, 5 or 18
Updated ยท Lumen Lab
All of the numbers are right; they measure different things. ๐จโ๐ฉโ๐ง is 1 character as people see it, 5 Unicode code points, 8 UTF-16 code units and 18 bytes in UTF-8.
The pieces inside one emoji
The family emoji is three emoji glued together with an invisible joiner (ZWJ).
๐จZWJ๐ฉZWJ๐ง
That is 5 code points. Written in UTF-16 they take 8 units, and in UTF-8 they take 18 bytes. On screen it fills one space, so to a reader it is 1 character.
Counted four ways
Counted with the โtaken apartโ panel on this site.
Text
Characters
Code points
UTF-16 units
UTF-8 bytes
a
1
1
1
1
รฉ (one code point)
1
1
1
2
รฉ (e plus a combining accent)
1
2
2
3
๐
1
1
2
4
๐๐ฝ (thumb plus skin tone)
1
2
4
8
๐บ๐ธ (U plus S)
1
2
4
8
๐จโ๐ฉโ๐ง
1
5
8
18
๐ด๓ ง๓ ข๓ ฅ๓ ฎ๓ ง๓ ฟ (flag of England)
1
7
14
28
Ad
What each number measures
Characters are what a reader would point at and call one letter or one emoji. Unicode calls them grapheme clusters, and the rules are in UAX #29. This is what the counters on this site report.
Code points are the numbered entries in Unicode. Combining accents, skin tones and joiners are code points of their own, which is how one visible character becomes several.
UTF-16 units are what JavaScript and several other languages count when you ask for the length of a string. A form that validates length that way sees ๐จโ๐ฉโ๐ง as 8.
Bytes are what gets stored or sent. UTF-8 uses 1 to 4 bytes per code point.
The same letter, two spellings
รฉ can be stored as one code point, or as a plain e followed by a combining accent. Both look identical. They differ in every column of the table except the first, and a search for one may not find the other. Software that cares usually normalises text to one form before counting. The twitter-text rules X uses do: either spelling counts as 1 there.
Which number does a limit use?
It depends on who set the limit, and the label rarely says. If an emoji uses up a limit faster than you expect, the form is counting units or bytes rather than characters. Two limits are documented well enough to work out exactly: X uses a weighted length, and SMS counts 7-bit slots or UTF-16 units.