I find myself wanting to block Reddit°hard word, etc°hard word. on the router°hard word level. There is nothing worth°hard word reading there.
----
the real way to get an "a-ha°hard word" moment from the machine, is with time-travel°hard word. 🔥 with time travel, the machine can return a better answer instantly°hard word!
the "head" of the response°hard word has to be noticeably°hard word ahead of where "committed°hard word" responses°hard word are. so, there is a possibility°hard word to jump backwards°hard word.
🔥 isn't this beam-search°hard word?
💡 ... maybe.
but, the idea is: you can assert°hard word a token°hard word ERROR PATH: RETREAT°hard word 32.
then°hard word, the "reason" for the message can be added as input°hard word.
it is the a-ha°hard word moment. but, implemented°hard word better than DeepSeek°hard word.
----
the reaction to DeepSeek°hard word has been, in my estimation°hard word, ridiculous°hard word.
I tried the 7b and 8b distilled°hard word models. And what I saw was a cheap parody°hard word of thought. Thought-processes°hard word that didn't make sense, and didn't actually reflect°hard word how the machine generated thoughts.
but, apparently, people like it.
Maybe the 400b model gives better answers? Or, maybe, people just see the shape of the answer and trust it more.
💡 if the goal°hard word of the machine is to solve°hard word industrial°hard word tasks, this is mostly already baked°hard word in to my estimations°hard word. but, for the goal°hard word of making consumers°hard word happier, there is clearly a factor°hard word I am not considering.
----
theory 1: people don't want to think the machine is smart°hard word; they want the machine to make them feel smart°hard word. both the appearance°hard word of struggling°hard word and the visible°hard word chain-of-thought°hard word (even if obviously flawed°hard word) contribute°hard word to this feeling.
theory 2: people don't know that the machine could already do 90% of this 12 months ago. they see a demo°hard word (or, more likely, hear about a demo°hard word) and, miraculously°hard word, now they know what will happen.
theory 3: we know that a light human touch guiding°hard word the machine's°hard word responses°hard word can improve accuracy°hard word substantially. and, that human touch can also be automated°hard word.
⚔️ well, actually, apparently very few other people knew that.
----
i'm going to stick with Theory 1 for today. that people like Deepseek°hard word (and feel it is better) because it makes them feel smart°hard word.
which ... is depressing°hard word. but, also, easily solvable°hard word.
the question is: what question could you pose°hard word that would lead somebody°hard word to come up with this answer on their own?
💡 it seems unlikely°hard word that 8B models can do this. but I assume°hard word the 600B models can.
----
people want the machine to make them feel smarter°hard word. 🔥 because people are self-centered°hard word, gullible°hard word, and insecure°hard word.
💡 they want it to behave°hard word in a way that I instinctively°hard word hate. they want the PT°hard word Barnum°hard word version°hard word of AI°hard word.
💬 give the people what they want!
----
this is probably one of the reasons why the default°hard word tone for every chatbot°hard word is obsequious°hard word. so much that's a great question! / you're absolutely right / let me know what else i can do to help.
----
💡 one can apply a politeness°hard word filter°hard word to the output°hard word of the machine. but the latency°hard word of such a system is already high.
⚔️ well, actually, it probably is just another layer or two.
----
Seen on social°hard word media°hard word: Anthropic°hard word is losing because they have rate limits! 💡 of course they have rate limits. the machine is not too cheap to meter, at least at the quality people expect.
----
a game of chess.
the idea of the attack worked in theory. and the attack worked in practice. but the actual°hard word attack did not work, in theory.
✨ chess can be a ritual°hard word. like the i-Ching°hard word.
----
can the machine participate°hard word in rituals°hard word?
----
there are two kinds of answers people want from the machine.
- Answers where people are willing to wait 10 minutes to definitely°hard word have the "right" answer.
- Entertainments°hard word. The various°hard word "instant°hard word chat-bots°hard word" are party tricks. A very good party trick. But, ultimately°hard word, a party trick.
🔥 perhaps testing is a third category°hard word.
Whereas°hard word: for many valuable°hard word use-cases°hard word, having a 5-minute latency°hard word to do it right, is not objectionable°hard word. 💡 the evocative questions, the "what do you mean by LONDON" and "can you talk more about LONDON", will be interactive°hard word. ⚙️ we do not have LONDON implemented°hard word here yet°hard word.