Posted by Alexander Power
Received: 2025-07-25 15:57:07
Channel: Cities - Project Journal
In reply to: lake andes, part 2 (View Chain)
Replies:
Keyboard shortcuts: E expand all, C collapse all, R restore default, Esc close open panels.
tasks for the next 24 hours:
- come up with a "cookbook" corpus ⚙️ consisting of recipes and descriptions of foods for wordfreq
- get a LLM-script to take a list of 2000 words and return the LONDON ⚙️ LONDON is a placeholder; it could be color or positive adjective or words like devil words.
- write a script to print the "top 2000 words by part-of-speech". ⚔️ well, actually, I already have these lists ... from the previous version of the database 💡 maybe "find the list" is more accurate
- write a script that will populate a few of the "sub-dictionaries" in the new format 💡 Countries, Nationalities, Numbers, and Colors will be the first 4, as they are fairly easy to check for completeness
💡 will the "Trakaido sub-dictionaries" be the canonical source-of-truth for what the GUIDs are? A flat-file is more cumbersome than a database for adding languages, linking to derivative forms, etc. But, it is easier for humans to read, and to put in Git repos.
tasks for the 48 hours after that:
- ensure the categories are stored in the Lemma table.
- re-assess the "GrammaticalForm" enum, because it doesn't work across languages. Maybe it needs to be "EnglishGrammaticalForm", "LithuanianGrammaticalForm", etc.
- generate "verb forms" - which requires some form of "aggregation" of WordToken entries
- generate all the "Level 1-5" entries from wordfreq 💡 right now, colors like orange are excluded because they are recent borrowings in Lithuanian. this is a very language-specific choice.
- consider how to handle "phrases" 💡 if you learn Malonu susipažinti before Malonu, it's a phrase and sentences