Kankan

Guides

Why a 161-character text costs two messages

Updated · Lumen Lab

A single SMS holds 160 characters of the GSM-7 alphabet. Once a text is split, every part gives up 7 of those slots to a header that tells the phone how to put the parts back together, so 161 characters travel as 153 + 8.

Where the breaks fall

Counted with the SMS counter on this site, using plain letters.

LengthPartsHow it splits
1601160
1612153 + 8
3062153 + 153
3073153 + 153 + 1
4604153 + 153 + 153 + 1

The 7 missing slots are the concatenation header defined in 3GPP TS 23.040. It is 6 bytes long, and 6 bytes need 7 of the 7-bit slots. A message that fits in one part does not carry it, which is why the first limit is 160 and every later one is 153.

One character can change the alphabet

GSM-7, defined in 3GPP TS 23.038, has room for basic Latin letters, digits, common punctuation and a short list of accented letters. If a text contains anything else, the whole message is sent as Unicode (UCS-2) instead, and the limits become 70 for a single message and 67 per part.

TextParts
160 letters1
160 letters and one 😀3
100 letters, é, 100 letters2
100 letters, ê, 100 letters3

The last two rows differ by one accent. é is in the GSM-7 alphabet; ê is not.

Ad

Characters that look harmless

These are outside GSM-7 and switch the whole text to Unicode:

  • Curly quotes and apostrophes: ’ “ ”. Phones and word processors often insert them for you, so “it’s” can cost more than “it's”.
  • The long dash and the single-character ellipsis (…).
  • The backtick (`) and the tab character.
  • Letters such as ç, á and ê. The capital Ç is in the alphabet, the small ç is not.
  • Every emoji. Most take 2 of the 70 units, so 69 letters and one emoji is 71 units and two parts.

These are fine: é, è, à, ñ, ü, ö.

Characters that take two slots

A few characters stay in GSM-7 but are sent as an escape code plus a character, so each uses 2 slots: [ ] { } ^ ~ | \ and €. A text of 80 curly brackets fills a message exactly; 81 make it two parts. A part never ends in the middle of one of these pairs, so the first part of that 81-bracket text carries 152 slots, not 153.

What this does not tell you

The part count is what the standard says the text needs. What you pay depends on your carrier or messaging provider, and some networks support national tables for Turkish, Spanish and Portuguese that change which letters fit. The counter does not model those.

Paste a text to see its parts and the characters that force Unicode.

Open the SMS counter

Keep reading

Ad