this post was submitted on 09 Jan 2025
176 points (98.4% liked)
Technology
60704 readers
7264 users here now
This is a most excellent place for technology news and articles.
Our Rules
- Follow the lemmy.world rules.
- Only tech related content.
- Be excellent to each another!
- Mod approved content bots can post up to 10 articles per day.
- Threads asking for personal tech support may be deleted.
- Politics threads may be removed.
- No memes allowed as posts, OK to post as comments.
- Only approved bots from the list below, to ask if your bot can be added please contact us.
- Check for duplicates before posting, duplicates may be removed
Approved Bots
founded 2 years ago
MODERATORS
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
Kinda shows there is a limit to how far you can get simply ingesting all text that exists. At some point, someone is going to need to curate perhaps billions of documents, which just based on volume will necessarily be done by people unqualified to really do so. And even if it were possible for a small group of people to curate such a data set, it would become an enormously political position to be in.
We did curation of existing knowledge for years, in the form of textbooks and reference works. This is just people thinking they can get the same benefits without the expense, and it'll come crashing down soon enough when people see that you need to handle concepts, not just surface words with a superficial autocomplete
Weird that they don't just...you know...copy that.
Even curation seems unlikely to fix the problem. I bet a new algorithm is required that allows LLMs to validate their response before it’s returned. Basically an “inner monologue” to avoid saying stupid things.
These models are so shit they need a translator. Hilarious.
I could use one of those...
validate against what? The "inner monologue" is the llm itself. It won't be any better than itself.
I swear to god, I feel like all of these LLM circlejerking shills have systematically forgotten one of the foundational points of computer science: garbage in, garbage out.