About Kankan
Kankan counts text. The same text gives different numbers depending on whether spaces are included, how a line break is counted, and whether you count characters or bytes. Kankan puts those rules side by side.
The name comes from 칸 (kan), one square on Korean manuscript paper: one square, one character.
The rules and where they come from
- Characters: grapheme clusters as defined in Unicode Standard Annex UAX #29, through the browser's Intl.Segmenter. Emoji sequences follow UTS #51.
- Words and sentences: words are runs of text between spaces, with a second figure that uses the UAX #29 word rules. Sentences use the UAX #29 sentence rules.
- Bytes: UTF-8 and UTF-16 as defined by the Unicode Standard. The Korean pages also count EUC-KR using the Windows code page 949 table.
- X weighted length: twitter-text 3.1.0, the open-source library X published for counting posts (checked 2026-10-09).
- SMS: the GSM 7-bit alphabet in 3GPP TS 23.038 and the concatenation rules in TS 23.040.
How the results are checked
The counting code is kept apart from the page and is covered by about 23,000 automated tests. The expected values were not produced by writing the same logic twice. They come from Python's standard codecs and Unicode libraries, from the twitter-text package, and from three independent SMS calculators.
Where your text goes
Nowhere. Text is counted inside your browser, and this site has no server that receives it. Ads are served by Google AdSense. The details are in the privacy policy.
Who makes it
Kankan is made by Lumen Lab. If you find a wrong number, write to woxocoso@gmail.com.