Counting words and characters online

The same text rarely produces the same total from one tool to the next. Here is what sits behind each figure.

You paste the same paragraph into Word, then into an online counter, then into a form field capped at 300 words. Three different totals, and an error message. None of the three is wrong: they are not counting the same thing. Once you understand what each one means by "word" and by "character", the problem goes away for good.

Why no two counters ever agree

There is no standard for counting. Every tool applies its own rules, and the discrepancies always turn up in the same places:

  • the scope taken into account: footnotes, headers, captions, text boxes, headings — included here, excluded there;
  • how hyphens and apostrophes are handled;
  • what happens to tokens with no letter in them: "%", "—", "…", a bullet number;
  • the unit of "character" in use, which is by far the trickiest point of all.

In a 500-word text, a dozen compound words — or a single footnote counted on one side and not the other — will stop the two totals from matching. Neither here nor there for a blog post; decisive when the brief sets a hard ceiling.

The character: three definitions living side by side

The word "character" means, depending on who is speaking, a UTF-16 code unit, a Unicode code point, or a grapheme — what the eye reads as a single sign. On plain Latin text the three coincide. The moment an emoji, or an accent stored as two separate characters, joins in, they part company.

A simple emoji such as ? is one code point but two UTF-16 units: a naive counter shows 2. A flag such as ?? is built from two code points, so four units: it shows 4. A "family" emoji, assembled with zero-width joiners, climbs to eleven. Convertu's word and character counter segments the text into graphemes: each of those signs counts as 1, exactly as it appears on screen.

The other trap is the invisible one: Unicode normalisation. The "é" in "café" exists in two forms, a precomposed character (NFC) or an "e" followed by a combining accent (NFD). Both look identical, but the second weighs two code points. Text exported from macOS or pulled out of certain PDFs frequently arrives in NFD, which inflates the total for no visible reason. A grapheme count is unaffected; a raw count is not.

The word: where the rules diverge

The most widespread convention, and the one Convertu applies, splits the text on whitespace and then keeps only the chunks holding at least one letter or digit. In practice that means:

  • "don't" and "mother-in-law" each count as a single word. Tools that break on the apostrophe or the hyphen turn them into two and three.
  • "50%" counts as one word, and so does "50 %" typed with a space: the percent sign on its own contains neither a letter nor a digit, so only the figure is counted.
  • A URL pasted into the body text counts as one word, however long it is.
  • A row of dashes, a string of dots or an orphaned bullet count for nothing at all.

Above all, remember this: when a brief is expressed in words — a 250-word abstract, a 500-word essay, translation billed by the source word — aim 3 to 5 per cent under the ceiling. That absorbs the gap between your counter and the one used by whoever is checking, and saves you hacking the text about at the last minute.

The character limits worth knowing

ContextUsual limitWhat catches people out
Title tagaround 60 charactersGoogle truncates on pixel width, not on the number of signs: a title in capitals fares worse than the same title in lower case.
Meta description150 to 160 charactersThe end is replaced by an ellipsis, often right in the middle of your best argument.
SMS160 in the GSM-7 alphabet, 70 otherwiseOne emoji, or a curly apostrophe pasted in from Word, tips the whole message into UCS-2 and cuts the capacity by more than half.
Google search ad30 per headline, 90 per descriptionSpaces count, and no formatting is allowed to claw back space.

For the first two rows there is no need to count by hand: the meta tag generator shows a live counter on the title and the description, and flags the overrun as you write.

Sentences, paragraphs and reading time

Sentence splitting relies on a full stop, an exclamation mark, a question mark or an ellipsis, provided it is followed by a space or by the end of the text. That refinement stops "3.14" being read as two sentences, but it offers no protection against abbreviations. British style at least spares you the full stop after "Mr" and "Dr"; "e.g.", "etc." and "Fig. 3" each create a phantom sentence all the same. On an official document stuffed with abbreviations, expect an overestimate of several units.

Paragraphs, meanwhile, are marked out by a blank line. Text pasted from an editor that separates its blocks with a single line break will therefore be seen as one long paragraph. When the figure looks absurdly low, that is almost always the explanation.

The reading time shown is based on 200 words per minute, an average silent reading speed for an adult on everyday prose. Two caveats. Dense technical writing is read appreciably more slowly. And, more to the point, reading aloud runs closer to 140 to 160 words per minute: for a video script or a talk, add roughly a third to the figure displayed.

The cases where the count stays wrong

The text is locked inside a PDF

You cannot count what you cannot select. If the PDF holds a text layer, a PDF to Word conversion hands back an editable document whose content copies normally — the file then passes through our servers, which erase it after an hour, so for anything confidential prefer selecting and copying the text straight out of the PDF. If the PDF is a scan — a photograph of a page, with no text underneath — no conversion will help: it needs optical character recognition first. The test takes two seconds: try to select a word in the document. If the cursor will not catch it, there is no text there to count.

These are entries, not words

For a list of email addresses, product references or keywords, the word count tells you nothing useful: what matters is the number of distinct lines. The duplicate remover shows the lines analysed, the unique lines and the duplicates stripped out, with a comparison you can set to ignore case and leading or trailing spaces.

The text is confidential

A contract, a job application, an exam script, an extract from a medical file: the counter runs entirely in the browser, and nothing you paste is sent to a server. That is true of 41 of the 60 tools in the catalogue. The other 19 — PDF, video, HEIC images — process the file server-side, with one free conversion per visitor per day and no sign-up; beyond that, a single subscription at €7 a year unlocks the lot, and uploaded files are deleted after an hour.

Related articles

20 online tools that replace installed software — and the one setting to know for each

Twenty utilities that sort out an unreadable HEIC, a PDF too large for a form or a 450 MB video, and the one setting to…

7 min read

Convertu vs Smallpdf vs iLovePDF: the 2026 comparison

All three make the same promises. Here are the four technical criteria that genuinely tell them apart, and the 30-secon…

6 min read

€7 a year: what the Convertu subscription actually pays for

Why 41 Convertu tools cost nothing to run, why the other 19 cost real money, and what your €7 a year actually pays for.

6 min read

← All articles