← Back to context

Comment by hn_throwaway_99

3 hours ago

Not the person you are responding to, but for me I really like to have a general understanding of how things work under the covers, but then have confidence that the contracts between me and those lower layers are rock solid, and that's simply not the case with LLMs.

Take software. In college I took computer architecture and digital logic design courses, and I thought they were immensely valuable in understanding how computers actually work and what software is actually doing. Sure, modern chip design is obviously several orders of magnitude more complex that what I studied, but I understand the basic concepts, and more importantly I have faith that chip designers do understand the nitty gritty details. Moving up the stack, I also had to build a rudimentary compiler in college. Again, modern compilers are a lot more complex, but I know that if something breaks, there is an identifiable bug either in my code or the compiler (or maybe even the chip). And importantly, while I may not have the skills to debug all the layers, when someone explains the bug to me, I can understand it in context (e.g. I'm not a chip designer but I thoroughly understand how Spectre is exploited and mitigated).

LLMs are nothing like that. Not even their builders understand the low level details. They're inherently stochastic systems, so people get slightly different results every time they're run. There is no "clean interface with a contract". And perhaps most importantly, many programmers are still expected to be responsible for their code, even when AI agents are generating so much of it that it's impossible to understand (or even read) it all. That's the thing that really stresses me out, when I'm responsible for a system but I don't really understand how it works.

To me the reasoning of understanding all the things you use in a rock solid way is hard to grasp.

You just typed some text and hit send. You trust that the combination of letters that form words and the combination of words that form sentences, arrive somewhere in a list of comments? But do you understand how those bytes (utf8 or utf16 or..) are transmitted (which endianness, how are they packaged, does it use sentinels or length codes), where do they end up (database, plain text), how is the text organised (alphabetical, by data, by points...)? In fact you are not sure about any of those. Still you typed and hit send. So how is that much different from writing a prompt and hit send?

  • So first of all, when LLMs first came on the scene, I did want to understand them better. I read the Attention is All You Need paper, I did Karpathy's transformer course, etc. And I think this basic understanding of LLMs and further reinforcement techniques makes me understand better how to use LLMs, e.g. what their limitations are, etc.

    But most importantly, there are very different qualifications in my head for things I just use and then get on with my day, and tools that are integral to things I am building. I take the bus but I don't really understand how a diesel engine works. But if I'm building it, I'm responsible for it. In fact, you brought up some networking issues in your comment. For a long time when I was a software engineer (primarily a web developer) I felt that my knowledge of network engineering was lacking, so I specifically took some courses in network engineering to better understand how my code interacted with the network. You also brought up text encodings. Once at my job I did a full, rabbit-hole deep dive into text encodings and localization because we had frequent bugs related to localization, and I still find a lot of developers misunderstand a bunch of the important details in localization (e.g. the difference between an encoding and a unicode code point).

    But sure, I don't need to understand every minutia of detail at the atomic detail. But I really like to know that I could, and in the meantime I have a enough of an understanding to build a complete conceptual model in my head.