Comment by scruple
2 hours ago
This feels like a sleight of hand to me. The hard part of evolving Scribe and Web Scrapbook was discovering that a browser extension manipulating a local SQLite database was _the only_ architecture that could reconcile local offline persistence with live DOM scraping across arbitrary catalogs of academic data.
An agent can synthesize existing solutions but (because I see this failure mode at work constantly) it can't synthesize an architecture to resolve the sorts of tensions that the person prompting it doesn't yet understand (not that that is stopping anyone). You can't prompt it to build something if the operational primitives required to solve the problem haven't been mapped.
"Build a tool based on Scribe and Web Scrapbook" in 2003 would've made a fragile PHP wrapper because that's what the existing landscape looked like.
Yes, exactly. And this is why I don't think that the LLMs can make significant process beyond what humans have done and published.
"But the math proofs," people will say. A lot of those seem to be spam-solving things with a huge swath of existing lemmas, and a some of these are being debunked and retracted.
Just today I was quizzing ChatGPT about a basic grammar question for a language that has huge training data but for which the grammar was not well documented. It kept giving me confidently wrong answers until I drilled and drilled it and then finally it found/gave back an explanation that perfectly fit a pattern given in one particular grammar, citing that as a source. It doesn't appear to have been able to figure out the inner structure on it's own. It appears only able to pattern match and put things together from what humans have already discovered and written.
Right, there's a difference between statistical interpolation and semantic induction. The whole point is that LLMs can't reason from first principles to drive missing rules. It keeps confidently feeding you approximations until it collides with some source that already mapped it.