today: https://spaceship.computer/greenland/model_summary.html
These are "proficiency°hard word" metrics°hard word. 🔥 although, every time I use the word "proficiency°hard word" I want to change it
They are simple tasks, currently: translate a word, choose a definition°hard word, choose an antonym°hard word, find the misspelled°hard word word. And, for the >4B models, as°hard word long as°hard word the model knows the language, it does fairly well. The 1B models do have some difficulties°hard word.
The timing data is interesting. It is, roughly, a linear°hard word relation to model size. The 9B models are about 4 times slower than the 1B models. Phi°hard word-4 (the largest°hard word model tested) is also very clearly the slowest model.
Some of the models I was looking at before (Granite, ExaONE°hard word, Hermes°hard word, Tulu°hard word, Mistral°hard word) did not°hard word make this round of tests. For Mistral°hard word, the 12B model is too old, and their newest release°hard word, at 24B, is too large. The others didn't°hard word distinguish°hard word themselves°hard word enough from similar Llama°hard word models to be worth°hard word my time (and hard-drive°hard word space).
remaining todo°hard word:
- standardize°hard word the logging°hard word of prompts°hard word and responses°hard word. the full text°hard word ⚙️ that is, including°hard word the system prompt°hard word should be stored°hard word.
- fix the benchmarks°hard word. some of the definitions°hard word are too similar. ⚙️ previously we had kingdom°hard word and realm°hard word as°hard word choices. now the closest°hard word is honest and sincere°hard word. some of the translations°hard word are still a bit rough. 💡 the translation°hard word of "beautiful" into French°hard word is beau°hard word/belle°hard word, the LLMs°hard word are very reasonably°hard word just returning "beau°hard word" as°hard word the translation°hard word
- fix the model warming. Just calling the "warm model" function°hard word correctly doesn't°hard word do enough warming.
- add additional°hard word tests. hopefully now it will take less than 1 hour to make new tests.
some of the suggestions regarding°hard word new tests:
Part of Speech Tagging°hard word - Present a sentence and ask the model to identify°hard word the part of speech (noun°hard word, verb°hard word, adjective°hard word, etc°hard word.) for a specific°hard word word.
Unit Conversion°hard word - Test ability°hard word to convert°hard word between simple units (kilometers to miles, pounds°hard word to kilograms).
Analogies°hard word - Simple analogies°hard word like "day is to night as°hard word hot is to ___".
Tense°hard word Transformation°hard word - Provide°hard word a sentence in one tense°hard word and ask the model to convert°hard word it to another tense°hard word.
Active°hard word/Passive°hard word Voice Conversion°hard word - Convert°hard word sentences between active°hard word and passive°hard word voice.