Three words, one rank, three different shapes

Show hard words

"Water", "number" and "large" sit almost side by side in our combined English frequency ranking, at #111, #117 and #124. If that were all you knew, you'd treat them as equally common words. But the combined rank is an average over 18 corpora, from cookbooks to Supreme Court opinions. Break it down by corpus and the three words behave very differently.

"Water" swings the most. In cookbooks it's the 15th most common word, at about 7,400 uses per million words. That puts it among words like "the" and "and". In Wikipedia's mathematics articles it falls to rank 3,265, at 23 per million: over 300 times rarer. It stays strong wherever the physical world is the subject (physical science, biology, and early modern science from Boyle to Osler). It fades in history, the arts and linguistics. How often you meet "water" depends almost entirely on what you're reading.

"Number" is lopsided the other way. It's the 27th most common word in math articles, and its next-best showing is in linguistics articles, where grammatical number (singular vs. plural) may give it a boost. It's also stronger than you'd expect in European Parliament debates, probably thanks to phrases like "a number of". In 19th- and 20th-century books, though, it falls below rank 880. Novelists apparently don't say "number" much. It's a word of expository writing.

"Large" is the steady one. Its best corpus (biology, rank 56) and its worst (Supreme Court opinions, rank 745) are only about 9× apart in frequency, against roughly 40× for "number" and 300× for "water". It turns up in cookbooks ("a large bowl"), geography, science and history alike. For a language learner, that makes "large" arguably the safest of the three to learn early: whatever you end up reading, you'll run into it.

💡 Claude's writing ability is still mediocre.

Rank alone hides dispersion. Words with nearly the same rank can have very different profiles: a steady word ("large"), a genre word ("number"), or a topic word ("water"). Corpus linguists measure this separately as dispersion (Juilland's D, or Gries' DP). For choosing what learners see first, a word that's moderately common everywhere is probably worth more than one that's very common in one place. The ratio between a word's best and worst corpus is a crude version of this metric, and it would be cheap to show next to the combined rank. Your new broadly-common-words report may already be heading this way.

The corpus mix shapes the ranking. About 12 of the 18 corpora are expository: Wikipedia, OpenStax and early-modern science. Only the 19th- and 20th-century books cover narrative writing, and nothing covers conversation. That helps "number" and probably pushes down words that dominate fiction and speech. The combined rank measures "common in informational English" more than "common in English".

A spike means topic, not frequency. When a word ranks far higher in one corpus than in the rest ("water" at #15 in cooking), the corpus is telling you its subject. You could flag these outliers automatically, which is the idea behind keyness in corpus linguistics.

Corpora count strings, not meanings. "Water" the noun and "water" the verb (which CEFR lists at B2) are counted as one token. "Large" has its frequency split 50/50 with "big" because of how the synonym links work. Any word with several parts of speech or senses has a combined number that is partly an artifact of these choices.