Comment by hn_throwaway_99
5 hours ago
Not the person you are responding to, but for me I really like to have a general understanding of how things work under the covers, but then have confidence that the contracts between me and those lower layers are rock solid, and that's simply not the case with LLMs.
Take software. In college I took computer architecture and digital logic design courses, and I thought they were immensely valuable in understanding how computers actually work and what software is actually doing. Sure, modern chip design is obviously several orders of magnitude more complex that what I studied, but I understand the basic concepts, and more importantly I have faith that chip designers do understand the nitty gritty details. Moving up the stack, I also had to build a rudimentary compiler in college. Again, modern compilers are a lot more complex, but I know that if something breaks, there is an identifiable bug either in my code or the compiler (or maybe even the chip). And importantly, while I may not have the skills to debug all the layers, when someone explains the bug to me, I can understand it in context (e.g. I'm not a chip designer but I thoroughly understand how Spectre is exploited and mitigated).
LLMs are nothing like that. Not even their builders understand the low level details. They're inherently stochastic systems, so people get slightly different results every time they're run. There is no "clean interface with a contract". And perhaps most importantly, many programmers are still expected to be responsible for their code, even when AI agents are generating so much of it that it's impossible to understand (or even read) it all. That's the thing that really stresses me out, when I'm responsible for a system but I don't really understand how it works.
To me the reasoning of understanding all the things you use in a rock solid way is hard to grasp.
You just typed some text and hit send. You trust that the combination of letters that form words and the combination of words that form sentences, arrive somewhere in a list of comments? But do you understand how those bytes (utf8 or utf16 or..) are transmitted (which endianness, how are they packaged, does it use sentinels or length codes), where do they end up (database, plain text), how is the text organised (alphabetical, by data, by points...)? In fact you are not sure about any of those. Still you typed and hit send. So how is that much different from writing a prompt and hit send?
Now consider the case where I send you a message (in a language that's not english) and it doesn't render as text on your screen. Reasonably, because I understand how network protocols work, I can assume that it's not a little/big endian mixup between our systems - the HTTP request would not have worked, then, as utf-8 and utf-16 are incompatible (unlike the weird utf-8 ascii mix). With HN, I can open up developer tools see what encoding it is sent as over-the-wire, and then I can ask you to look at the encoding on your end. I can say "the vibecoded browser you use is trying to interpret this as utf-16 but hackernews only sends utf-8". I can dial down the root cause of the problem and then address it.
Contrast that to a situation where I ask an LLM to use a tool and it can only do it like 98% of the time. How do you root cause that. How do you fix it such that that doesn't happen again?
Also consider: occasionally your support agent gives people free plane tickets.
Yes, we've made a lot of progress towards LLMs being fairly good, and adversarial systems and such do work, but it is asymptotic approach to 100%, never actually 100%
So first of all, when LLMs first came on the scene, I did want to understand them better. I read the Attention is All You Need paper, I did Karpathy's transformer course, etc. And I think this basic understanding of LLMs and further reinforcement techniques makes me understand better how to use LLMs, e.g. what their limitations are, etc.
But most importantly, there are very different qualifications in my head for things I just use and then get on with my day, and tools that are integral to things I am building. I take the bus but I don't really understand how a diesel engine works. But if I'm building it, I'm responsible for it. In fact, you brought up some networking issues in your comment. For a long time when I was a software engineer (primarily a web developer) I felt that my knowledge of network engineering was lacking, so I specifically took some courses in network engineering to better understand how my code interacted with the network. You also brought up text encodings. Once at my job I did a full, rabbit-hole deep dive into text encodings and localization because we had frequent bugs related to localization, and I still find a lot of developers misunderstand a bunch of the important details in localization (e.g. the difference between an encoding and a unicode code point).
But sure, I don't need to understand every minutia of detail at the atomic detail. But I really like to know that I could, and in the meantime I have a enough of an understanding to build a complete conceptual model in my head.