I bet you can make a cheap “AI” chip for using the trained model, it seems the training is the ruinous thing today.
Remember when they predicted heavier than air flight to be centuries in the future just the week before the wright brothers flew? Me neither I wasn’t born then, but it’s an interesting example IMO.
The difference here is that a cheap AI chip won’t fix the fundamental software problems with LLMs. We might reach a point where they can produce output faster, but as long as what’s actually going on is probabilistic next-token prediction in a static vector database, that just means faster mistakes as well.
There’s an odd psychosis going around where people become convinced that actual AGI can be derived from this technology. People who should know better just shut their brains off when it comes to token prediction, because they’ve had very compelling “conversations” with the predictor. They forget that the actual model is static, has no internal state, and doesn’t even “remember” what you’ve said to it.
What it has is a context window, and your entire conversational history - both what you’ve said and how it has responded - gets shoved into that window when you interact with it. (Or depending on the chatbot harness, saved in “memory” files that it can retrieve when the context contents indicate that would be useful.)
That’s why the bots seem so weirdly forgetful one moment and like they’ve got photographic memories the next. Stuff that is in the context window and has its “attention” will influence the tokens it produces, but whether or not the right things are in the context window and it’s including them in the token prediction is a crapshoot.
I know right. “AGI” !!! Nah that won’t happen tomorrow Kevin.
But with thousands of the smartest engineers and researchers working on it, we might get a smarter AI. I mean the human brain thinks in similar ways. +a lot of other stuff of course, but maybe that stuff can be emulated, simulated, for “the next step” forward (still no agi lol).
I bet you can make a cheap “AI” chip for using the trained model, it seems the training is the ruinous thing today.
Remember when they predicted heavier than air flight to be centuries in the future just the week before the wright brothers flew? Me neither I wasn’t born then, but it’s an interesting example IMO.
The difference here is that a cheap AI chip won’t fix the fundamental software problems with LLMs. We might reach a point where they can produce output faster, but as long as what’s actually going on is probabilistic next-token prediction in a static vector database, that just means faster mistakes as well.
There’s an odd psychosis going around where people become convinced that actual AGI can be derived from this technology. People who should know better just shut their brains off when it comes to token prediction, because they’ve had very compelling “conversations” with the predictor. They forget that the actual model is static, has no internal state, and doesn’t even “remember” what you’ve said to it.
What it has is a context window, and your entire conversational history - both what you’ve said and how it has responded - gets shoved into that window when you interact with it. (Or depending on the chatbot harness, saved in “memory” files that it can retrieve when the context contents indicate that would be useful.)
That’s why the bots seem so weirdly forgetful one moment and like they’ve got photographic memories the next. Stuff that is in the context window and has its “attention” will influence the tokens it produces, but whether or not the right things are in the context window and it’s including them in the token prediction is a crapshoot.
I know right. “AGI” !!! Nah that won’t happen tomorrow Kevin.
But with thousands of the smartest engineers and researchers working on it, we might get a smarter AI. I mean the human brain thinks in similar ways. +a lot of other stuff of course, but maybe that stuff can be emulated, simulated, for “the next step” forward (still no agi lol).
Interesting times lie ahead.