Three words, one rank, three different shapes

Hide hard words

"Water", "number" and "large" sit almost side by side in our combined°hard word English°hard word frequency°hard word ranking°hard word, at #111, #117 and #124. If that were all you knew, you'd treat°hard word them as equally common words. But the combined°hard word rank°hard word is an average°hard word over 18 corpora°hard word, from cookbooks°hard word to Supreme°hard word Court opinions. Break it down by corpus°hard word and the three words behave°hard word very differently.

"Water" swings°hard word the most. In cookbooks°hard word it's the 15th°hard word most common word, at about 7,400 uses per°hard word million°hard word words. That puts it among words like "the" and "and". In Wikipedia's°hard word mathematics articles°hard word it falls to rank°hard word 3,265, at 23 per°hard word million°hard word: over 300 times rarer°hard word. It stays strong wherever°hard word the physical world°hard word is the subject (physical science, biology, and early modern°hard word science from Boyle°hard word to Osler°hard word). It fades°hard word in history, the arts and linguistics°hard word. How often you meet "water" depends°hard word almost entirely on what you're reading.

"Number" is lopsided°hard word the other way. It's the 27th°hard word most common word in math articles°hard word, and its next-best°hard word showing is in linguistics°hard word articles°hard word, where grammatical°hard word number (singular°hard word vs°hard word. plural°hard word) may give it a boost°hard word. It's also stronger than you'd expect in European°hard word Parliament debates°hard word, probably thanks to phrases°hard word like "a number of". In 19th°hard word- and 20th-century°hard word books, though, it falls below rank°hard word 880. Novelists°hard word apparently don't say "number" much. It's a word of expository°hard word writing.

"Large" is the steady°hard word one. Its best corpus°hard word (biology, rank°hard word 56) and its worst (Supreme°hard word Court opinions, rank°hard word 745) are only about 9× apart in frequency°hard word, against roughly 40× for "number" and 300× for "water". It turns up in cookbooks°hard word ("a large bowl"), geography, science and history alike°hard word. For a language learner, that makes "large" arguably°hard word the safest of the three to learn early: whatever°hard word you end up reading, you'll run into it.

💡 Claude's°hard word writing ability°hard word is still mediocre°hard word.

Rank°hard word alone hides dispersion°hard word. Words with nearly the same rank°hard word can have very different profiles°hard word: a steady°hard word word ("large"), a genre°hard word word ("number"), or a topic°hard word word ("water"). Corpus°hard word linguists°hard word measure this separately as dispersion°hard word (Juilland's°hard word D, or Gries'°hard word DP°hard word). For choosing what learners°hard word see first, a word that's moderately°hard word common everywhere is probably worth°hard word more than one that's very common in one place. The ratio°hard word between a word's°hard word best and worst corpus°hard word is a crude°hard word version°hard word of this metric°hard word, and it would be cheap to show next to the combined°hard word rank°hard word. Your new broadly-common-words°hard word report may already be heading this way.

The corpus°hard word mix°hard word shapes the ranking°hard word. About 12 of the 18 corpora°hard word are expository°hard word: Wikipedia°hard word, OpenStax°hard word and early-modern°hard word science. Only the 19th°hard word- and 20th-century°hard word books cover narrative°hard word writing, and nothing covers conversation. That helps "number" and probably pushes down words that dominate°hard word fiction and speech. The combined°hard word rank°hard word measures "common in informational°hard word English°hard word" more than "common in English°hard word".

A spike°hard word means topic°hard word, not frequency°hard word. When a word ranks°hard word far higher in one corpus°hard word than in the rest ("water" at #15 in cooking), the corpus°hard word is telling you its subject. You could flag these outliers°hard word automatically°hard word, which is the idea behind keyness°hard word in corpus°hard word linguistics°hard word.

Corpora°hard word count strings, not meanings. "Water" the noun°hard word and "water" the verb°hard word (which CEFR°hard word lists at B2) are counted as one token°hard word. "Large" has its frequency°hard word split°hard word 50/50 with "big" because of how the synonym°hard word links°hard word work. Any word with several parts of speech or senses has a combined°hard word number that is partly an artifact°hard word of these choices.