so far today: running the "proficiency°hard word" benchmarks°hard word against GPT°hard word-4-1 and Gemini°hard word-2.5-flash.
----
The headline°hard word: Google's°hard word cheap model can count letters. Gemini°hard word was substantially slower than both OpenAI°hard word and Anthropic°hard word (but, perhaps, that can vary°hard word day-to-day°hard word). But it got 96% on the infamous°hard word count how many "R"s in strawberry metric°hard word, and none of the similarly-priced°hard word models got above 70%. 💡 the only metric°hard word it did "bad" on was the IPA°hard word one, and that is because the response°hard word normalization°hard word code°hard word is broken
----
for pricing°hard word ⚙️ all prices per°hard word million°hard word tokens°hard word:
GPT°hard word-4-1-nano°hard word: 10c IN, 40c OUT
GPT°hard word-4-1-mini°hard word: 40c IN, 160c OUT
GPT°hard word-4o-mini°hard word: 30c IN, 120c OUT
Gemini°hard word-2.5-flash: 15c IN, 60c OUT
Claude°hard word-3-5-haiku°hard word: 80c IN, 400c OUT
⚙️ Most of these have (or will have) "cache°hard word" discounts°hard word of 50-90% for repeated°hard word queries°hard word with the same long context°hard word.
💡 Claude°hard word is both the most expensive°hard word at this tier°hard word, and the lowest-performing°hard word. And the least-recently°hard word updated°hard word.
🔥 presumably°hard word, they will have a new model at half the price, next week.