{"chain":[{"channel":"cities","content":"<green> <<< *Lake Andes* is located in southern South Dakota, part of the Yankton Sioux reservation. >>>\r\n\r\nthe goal for today: merge the \"Trakaido wordlist\" and the \"Greenland wordlist\".\r\n\r\n----\r\n\r\nThe Trakaido wordlist is 720 non-verbs (nouns, adjectives, and \"grammatical words\" like << with >> or << how >>), in various categories. (<red> there are also verbs, but in conjugated forms, not suitable for a \"merge\" yet.)\r\n\r\nThe Greenland wordlist is 4000-10000 << WordToken >>s, from word frequency lists.  These do not clarify specific meaning, or whether it is part of a multi-word phrase.\r\n\r\n----\r\n\r\nThe merged wordlist will probably be << lemma >> based, with a list of \"derivative forms\" for each. (<red> the issue that \"run\" and \"running\" are separate lemmas when they are different parts of speech ... I will call it something else if need be.)\r\n\r\n--MORE--\r\n\r\nThe current outputs, for Trakaido:\r\n\r\n<<< additional_foods.py:\r\nN16011 = {\r\n  'guid': 'N16011',\r\n  'english': 'ketchup',\r\n  'lithuanian': 'ke\u010dupas',\r\n  'alternatives': {\r\n    'english': [],\r\n    'lithuanian': []\r\n  },\r\n  'metadata': {\r\n    'difficulty_level': None,\r\n    'frequency_rank': None,\r\n    'tags': [],\r\n    'notes': ''\r\n  }\r\n} >>>\r\n\r\nand, for greenland:\r\n\r\n<<< Word: ketchup\r\nRank: None\r\n\r\nDefinitions:\r\n  [1] A thick, sweet, and tangy sauce made primarily from tomatoes, often used as a condiment for foods like fries, burgers, and hot dogs.\r\n    Confidence: 1.00\r\n    Lemma: ketchup\r\n    Part of speech: noun\r\n    Grammatical form: noun/singular\r\n    Subtype: food_drink\r\n      IPA: /\u02c8k\u025bt\u0283\u028cp/\r\n      Phonetic Pronunciation: KEH-chup\r\n    Chinese: \u756a\u8304\u9171\r\n    French: ketchup\r\n    Korean: \ucf00\ucca9\r\n    Swahili: ketchup\r\n    Lithuanian: ketchup\r\n    Vietnamese: t\u01b0\u01a1ng c\u00e0\r\n    Notes: This is the most common and primary meaning of 'ketchup'.\r\n    Examples:\r\n      - I like to put ketchup on my fries.\r\n      - Could you pass the ketchup, please?\r\n      - Ketchup is a popular condiment worldwide.\r\n>>>\r\n(<orange> the fact that we used a cheap model, asked for multiple definitions, and got << To apply ketchup to food. >>, as in << She will ketchup the fries after they arrive. >> ... is not relevant. ) (<red> the fact that it mistranslated ketchup into Lithuanian is more of an issue.)","created_at":"2025-07-24T17:04:20.900618","id":635,"is_target":false,"parent_id":null,"processed_content":"<div class=\"mlq color-green\"><button type=\"button\" class=\"mlq-collapse\" aria-label=\"Toggle visibility\"><span class=\"mlq-collapse-icon\">\u2699\ufe0f</span></button><div class=\"mlq-content\"><p> <em>Lake Andes</em> is located in southern South Dakota, part of the Yankton Sioux reservation. </p></div></div>\n<p>the goal for today: merge the \"Trakaido wordlist\" and the \"Greenland wordlist\".\r</p>\n<hr class=\"section-break\" />\n<p>The Trakaido wordlist is 720 non-verbs (nouns, adjectives, and \"grammatical words\" like <span class=\"literal-text\">with</span> or <span class=\"literal-text\">how</span>), in various categories. <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> there are also verbs, but in conjugated forms, not suitable for a \"merge\" yet.</span></span>\r</p>\n<p>The Greenland wordlist is 4000-10000 <span class=\"literal-text\">WordToken</span>s, from word frequency lists.  These do not clarify specific meaning, or whether it is part of a multi-word phrase.\r</p>\n<hr class=\"section-break\" />\n<p>The merged wordlist will probably be <span class=\"literal-text\">lemma</span> based, with a list of \"derivative forms\" for each. <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> the issue that \"run\" and \"running\" are separate lemmas when they are different parts of speech ... I will call it something else if need be.</span></span>\r</p>\n<div class=\"content-sigil\" aria-label=\"Extended content begins here\">&#9135;&#9135;&#9135;&#9135;&#9135;</div>\n<p>The current outputs, for Trakaido:\r</p>\n<div class=\"mlq\"><button type=\"button\" class=\"mlq-collapse\" aria-label=\"Toggle visibility\"><span class=\"mlq-collapse-icon\">-</span></button><div class=\"mlq-content\"><p> additional_foods.py:\r</p>\n<p>N16011 = {\r</p>\n<p>  'guid': 'N16011',\r</p>\n<p>  'english': 'ketchup',\r</p>\n<p>  'lithuanian': 'ke\u010dupas',\r</p>\n<p>  'alternatives': {\r</p>\n<p>    'english': [],\r</p>\n<p>    'lithuanian': []\r</p>\n<p>  },\r</p>\n<p>  'metadata': {\r</p>\n<p>    'difficulty_level': None,\r</p>\n<p>    'frequency_rank': None,\r</p>\n<p>    'tags': [],\r</p>\n<p>    'notes': ''\r</p>\n<p>  }\r</p>\n<p>} </p></div></div>\n<p>and, for greenland:\r</p>\n<div class=\"mlq\"><button type=\"button\" class=\"mlq-collapse\" aria-label=\"Toggle visibility\"><span class=\"mlq-collapse-icon\">-</span></button><div class=\"mlq-content\"><p> Word: ketchup\r</p>\n<p>Rank: None\r</p>\n<p>Definitions:\r</p>\n<p>  [1] A thick, sweet, and tangy sauce made primarily from tomatoes, often used as a condiment for foods like fries, burgers, and hot dogs.\r</p>\n<p>    Confidence: 1.00\r</p>\n<p>    Lemma: ketchup\r</p>\n<p>    Part of speech: noun\r</p>\n<p>    Grammatical form: noun/singular\r</p>\n<p>    Subtype: food_drink\r</p>\n<p>      IPA: /\u02c8k\u025bt\u0283\u028cp/\r</p>\n<p>      Phonetic Pronunciation: KEH-chup\r</p>\n<p>    Chinese: <span class=\"annotated-chinese\" data-pinyin=\"F\u0100N Q\u00cdE J\u00ccANG\" data-definition=\"ketchup\">\u756a\u8304\u9171</span>\r</p>\n<p>    French: ketchup\r</p>\n<p>    Korean: \ucf00\ucca9\r</p>\n<p>    Swahili: ketchup\r</p>\n<p>    Lithuanian: ketchup\r</p>\n<p>    Vietnamese: t\u01b0\u01a1ng c\u00e0\r</p>\n<p>    Notes: This is the most common and primary meaning of 'ketchup'.\r</p>\n<p>    Examples:\r</p>\n<p>      - I like to put ketchup on my fries.\r</p>\n<p>      - Could you pass the ketchup, please?\r</p>\n<p>      - Ketchup is a popular condiment worldwide.\r</p></div></div>\n<p><span class=\"colorblock color-orange\"><span class=\"sigil\">\u2694\ufe0f</span><span class=\"colortext-content\"> the fact that we used a cheap model, asked for multiple definitions, and got <span class=\"literal-text\">To apply ketchup to food.</span>, as in <span class=\"literal-text\">She will ketchup the fries after they arrive.</span> ... is not relevant. </span></span> <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> the fact that it mistranslated ketchup into Lithuanian is more of an issue.</span></span></p>","subject":"lake andes"},{"channel":"cities","content":"In Lithuanian, the term for \"Lithuanian language\" is << Lietuvi\u0173 kalba >>, for \"Lithuanian man\" is << lietuvis >>, and for \"Lithuanian woman\" is << lietuv\u0117 >>.\r\n\r\nThese are all the same word in English, << Lithuanian >>.  Here, we have separate files (<xantham> or \"grammatical form\") for \"nationality\" and \"language\".\r\n\r\n<red> with 200 nationalities and 400 languages, it is fine to have a \"special case\" for each.  I am expecting 100-200 \"special cases\", including 10-20 that are catch-alls.\r\n\r\n----\r\n\r\nThe thought of having \"Derivative Forms\" shared between languages is, unfortunately, impossible.\r\n\r\nSo, the \"lemma\" form (infinitive, etc.) will have to be translated.  This, still, is difficult.  Is it << la piscine >> or just << piscine >>?  Or, for that matter, \"pool\" or \"swimming pool\" (or << pool (swimming) >>).\r\n\r\nBut, then, there will be separate derivative forms for \"walks\", \"walked\", etc., all of which are mono-lingual.  For << marchons >>, << marchez >>, etc. it will be a separate set.\r\n\r\n----\r\n\r\nFor (conjunctions, prepositions, etc.) I want to still call them \"grammatical words\".\r\n\r\nBecause a word-for-word \"translation\" is too perilous to attempt.\r\n\r\n--MORE--\r\n\r\nIn practical terms, for verbs, Greenland output will have a pythondict like:\r\n\r\n<<<\r\nV0023 = {\r\n\"guid\": \"V0023\",\r\n\"english\": \"to eat\",\r\n\"lithuanian\": \"valgyti\",\r\n\"english_forms\" = {\"first-person-singular-present\": \"I eat\", ...}\r\n\"lithuanian_forms\" = {\"first-person-singular-present\": \"A\u0161 valgau\", ...} ...\r\n>>>\r\n\r\nThis is similar enough to the current Trakaido format:\r\n\r\n<<< \r\n\"valgyti\": {\r\n    \"english\": \"to eat\",\r\n    \"present_tense\": {\r\n      \"1s\": {\"english\": \"I eat\", \"lithuanian\": \"a\u0161 valgau\"},\r\n      \"2s\": {\"english\": \"you(s.) eat\", \"lithuanian\": \"tu valgai\"},\r\n      \"3s-m\": {\"english\": \"he eats\", \"lithuanian\": \"jis valgo\"}, ...\r\n>>>\r\n\r\nFor both, it will request \"valgyti / first-person-singular-present\" as a \"flashcard\".\r\n\r\n----\r\n\r\nThe new approach will simplify some of the \"more animals\" style groups.\r\n\r\nAll the animals will be in a single \"dictionary\" file.  The first 12 will be in \"Animals 1\", the next 18 in \"Animals 2\", etc.\r\n\r\n<red> this will *probably* come with a reworking of the \"corpus\" system.  We mostly want \"levels\" now anyway ... and \"decoy sets\" can be configured separately.","created_at":"2025-07-24T21:52:37.965355","id":636,"is_target":true,"parent_id":635,"processed_content":"<p>In Lithuanian, the term for \"Lithuanian language\" is <span class=\"literal-text\">Lietuvi\u0173 kalba</span>, for \"Lithuanian man\" is <span class=\"literal-text\">lietuvis</span>, and for \"Lithuanian woman\" is <span class=\"literal-text\">lietuv\u0117</span>.\r</p>\n<p>These are all the same word in English, <span class=\"literal-text\">Lithuanian</span>.  Here, we have separate files <span class=\"colorblock color-xantham\"><span class=\"sigil\">\ud83d\udd25</span><span class=\"colortext-content\"> or \"grammatical form\"</span></span> for \"nationality\" and \"language\".\r</p>\n<p><span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> with 200 nationalities and 400 languages, it is fine to have a \"special case\" for each.  I am expecting 100-200 \"special cases\", including 10-20 that are catch-alls.\r</span></span></p>\n<hr class=\"section-break\" />\n<p>The thought of having \"Derivative Forms\" shared between languages is, unfortunately, impossible.\r</p>\n<p>So, the \"lemma\" form (infinitive, etc.) will have to be translated.  This, still, is difficult.  Is it <span class=\"literal-text\">la piscine</span> or just <span class=\"literal-text\">piscine</span>?  Or, for that matter, \"pool\" or \"swimming pool\" (or <span class=\"literal-text\">pool (swimming)</span>).\r</p>\n<p>But, then, there will be separate derivative forms for \"walks\", \"walked\", etc., all of which are mono-lingual.  For <span class=\"literal-text\">marchons</span>, <span class=\"literal-text\">marchez</span>, etc. it will be a separate set.\r</p>\n<hr class=\"section-break\" />\n<p>For (conjunctions, prepositions, etc.) I want to still call them \"grammatical words\".\r</p>\n<p>Because a word-for-word \"translation\" is too perilous to attempt.\r</p>\n<div class=\"content-sigil\" aria-label=\"Extended content begins here\">&#9135;&#9135;&#9135;&#9135;&#9135;</div>\n<p>In practical terms, for verbs, Greenland output will have a pythondict like:\r</p>\n<div class=\"mlq\"><button type=\"button\" class=\"mlq-collapse\" aria-label=\"Toggle visibility\"><span class=\"mlq-collapse-icon\">-</span></button><div class=\"mlq-content\"><p>V0023 = {\r</p>\n<p>\"guid\": \"V0023\",\r</p>\n<p>\"english\": \"to eat\",\r</p>\n<p>\"lithuanian\": \"valgyti\",\r</p>\n<p>\"english_forms\" = {\"first-person-singular-present\": \"I eat\", ...}\r</p>\n<p>\"lithuanian_forms\" = {\"first-person-singular-present\": \"A\u0161 valgau\", ...} ...\r</p></div></div>\n<p>This is similar enough to the current Trakaido format:\r</p>\n<div class=\"mlq\"><button type=\"button\" class=\"mlq-collapse\" aria-label=\"Toggle visibility\"><span class=\"mlq-collapse-icon\">-</span></button><div class=\"mlq-content\"><p>\"valgyti\": {\r</p>\n<p>    \"english\": \"to eat\",\r</p>\n<p>    \"present_tense\": {\r</p>\n<p>      \"1s\": {\"english\": \"I eat\", \"lithuanian\": \"a\u0161 valgau\"},\r</p>\n<p>      \"2s\": {\"english\": \"you(s.) eat\", \"lithuanian\": \"tu valgai\"},\r</p>\n<p>      \"3s-m\": {\"english\": \"he eats\", \"lithuanian\": \"jis valgo\"}, ...\r</p></div></div>\n<p>For both, it will request \"valgyti / first-person-singular-present\" as a \"flashcard\".\r</p>\n<hr class=\"section-break\" />\n<p>The new approach will simplify some of the \"more animals\" style groups.\r</p>\n<p>All the animals will be in a single \"dictionary\" file.  The first 12 will be in \"Animals 1\", the next 18 in \"Animals 2\", etc.\r</p>\n<p><span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> this will <em>probably</em> come with a reworking of the \"corpus\" system.  We mostly want \"levels\" now anyway ... and \"decoy sets\" can be configured separately.</span></span></p>","subject":"lake andes, part 2"},{"channel":"cities","content":"tasks for the next 24 hours:\r\n\r\n# come up with a \"cookbook\" corpus (<green> consisting of recipes and descriptions of foods) for wordfreq\r\n# get a LLM-script to take a list of 2000 words and return the LONDON (<green> LONDON is a placeholder; it could be << color >> or << positive adjective >> or << words like devil >>) words.\r\n# write a script to print the \"top 2000 words by part-of-speech\". (<orange> well, actually, I already *have* these lists ... from the previous version of the database) (<red> maybe \"find the list\" is more accurate)\r\n# write a script that will populate a few of the \"sub-dictionaries\" in the new format (<red> Countries, Nationalities, Numbers, and Colors will be the first 4, as they are fairly easy to check for completeness)\r\n\r\n(<red> will the \"Trakaido sub-dictionaries\" be the canonical source-of-truth for what the GUIDs are?  A flat-file is more cumbersome than a database for adding languages, linking to << derivative forms >>, etc.  But, it is easier for humans to read, and to put in Git repos.)\r\n\r\n----\r\n\r\ntasks for the 48 hours after that:\r\n# ensure the categories are stored in the << Lemma >> table.\r\n# re-assess the \"GrammaticalForm\" enum, because it doesn't work across languages.  Maybe it needs to be \"EnglishGrammaticalForm\", \"LithuanianGrammaticalForm\", etc.\r\n# generate \"verb forms\" - which requires some form of \"aggregation\" of WordToken entries\r\n# generate all the \"Level 1-5\" entries from wordfreq (<red> right now, colors like << orange >> are excluded because they are recent borrowings in Lithuanian.  this is a very language-specific choice.)\r\n# consider how to handle \"phrases\" (<red> if you learn << Malonu susipa\u017einti >> before << Malonu >>, it's a phrase) and sentences","created_at":"2025-07-25T15:57:07.455679","id":638,"is_target":false,"parent_id":636,"processed_content":"<p>tasks for the next 24 hours:\r</p>\n<ul>\n<li class=\"number-list\"> come up with a \"cookbook\" corpus <span class=\"colorblock color-green\"><span class=\"sigil\">\u2699\ufe0f</span><span class=\"colortext-content\"> consisting of recipes and descriptions of foods</span></span> for wordfreq\r</li>\n<li class=\"number-list\"> get a LLM-script to take a list of 2000 words and return the LONDON <span class=\"colorblock color-green\"><span class=\"sigil\">\u2699\ufe0f</span><span class=\"colortext-content\"> LONDON is a placeholder; it could be <span class=\"literal-text\">color</span> or <span class=\"literal-text\">positive adjective</span> or <span class=\"literal-text\">words like devil</span></span></span> words.\r</li>\n<li class=\"number-list\"> write a script to print the \"top 2000 words by part-of-speech\". <span class=\"colorblock color-orange\"><span class=\"sigil\">\u2694\ufe0f</span><span class=\"colortext-content\"> well, actually, I already <em>have</em> these lists ... from the previous version of the database</span></span> <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> maybe \"find the list\" is more accurate</span></span>\r</li>\n<li class=\"number-list\"> write a script that will populate a few of the \"sub-dictionaries\" in the new format <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> Countries, Nationalities, Numbers, and Colors will be the first 4, as they are fairly easy to check for completeness</span></span>\r</li>\n</ul>\n<p><span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> will the \"Trakaido sub-dictionaries\" be the canonical source-of-truth for what the GUIDs are?  A flat-file is more cumbersome than a database for adding languages, linking to <span class=\"literal-text\">derivative forms</span>, etc.  But, it is easier for humans to read, and to put in Git repos.</span></span>\r</p>\n<hr class=\"section-break\" />\n<p>tasks for the 48 hours after that:\r</p>\n<ul>\n<li class=\"number-list\"> ensure the categories are stored in the <span class=\"literal-text\">Lemma</span> table.\r</li>\n<li class=\"number-list\"> re-assess the \"GrammaticalForm\" enum, because it doesn't work across languages.  Maybe it needs to be \"EnglishGrammaticalForm\", \"LithuanianGrammaticalForm\", etc.\r</li>\n<li class=\"number-list\"> generate \"verb forms\" - which requires some form of \"aggregation\" of WordToken entries\r</li>\n<li class=\"number-list\"> generate all the \"Level 1-5\" entries from wordfreq <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> right now, colors like <span class=\"literal-text\">orange</span> are excluded because they are recent borrowings in Lithuanian.  this is a very language-specific choice.</span></span>\r</li>\n<li class=\"number-list\"> consider how to handle \"phrases\" <span class=\"colorblock color-red\"><span class=\"sigil\">\ud83d\udca1</span><span class=\"colortext-content\"> if you learn <span class=\"literal-text\">Malonu susipa\u017einti</span> before <span class=\"literal-text\">Malonu</span>, it's a phrase</span></span> and sentences</li>\n</ul>","subject":"lake andes, part 3"}]}
