the project has been at the liminal°hard word point between "interesting" and "we already have dictionaries°hard word".
some things are easier°hard word for the LLM°hard word than others. for "get an IPA°hard word pronunciation°hard word", I eventually°hard word determined°hard word that the best choice was to give up and download°hard word a flat-text°hard word file°hard word. (and, maybe, get the machine to deal with the ambiguous°hard word cases)
----
the "merge°hard word various°hard word word-frequency°hard word lists into one list" code°hard word, after a few rounds of telling the machine it was wrong, now works fine.
💡 the top-level°hard word differences°hard word are predictable°hard word. words like you are less common on Wikipedia°hard word.
⚙️ because of some bug°hard word, words like "vernacular°hard word" and "justification°hard word" are showing up in the top 250. also, not°hard word all the lists manage°hard word contractions°hard word correctly, giving°hard word "words" like isn°hard word and doesn°hard word ('t°hard word).
----
perhaps the next task is "generate an annotated°hard word version°hard word of text°hard word, where the text°hard word is colored based°hard word on the word frequency°hard word".
💡 and, possibly, the uncommon°hard word words get Chinese°hard word translations°hard word added.
🔥 that will certainly help the English°hard word monoglot°hard word who is confused!
----
the variance°hard word in the word frequencies°hard word (mostly) says something about the "cultural°hard word loading" of the words.
most of the words that are more common in 19th°hard word century books are low-cultural-loading°hard word. 💡 perhaps the prevalence°hard word of gentleman°hard word is cultural°hard word. but words like rain and hat are more frequent just because other words (geometry°hard word, organic°hard word) are less frequent.