← Back to context

Comment by hn_throwaway_99

3 hours ago

So first of all, when LLMs first came on the scene, I did want to understand them better. I read the Attention is All You Need paper, I did Karpathy's transformer course, etc. And I think this basic understanding of LLMs and further reinforcement techniques makes me understand better how to use LLMs, e.g. what their limitations are, etc.

But most importantly, there are very different qualifications in my head for things I just use and then get on with my day, and tools that are integral to things I am building. I take the bus but I don't really understand how a diesel engine works. But if I'm building it, I'm responsible for it. In fact, you brought up some networking issues in your comment. For a long time when I was a software engineer (primarily a web developer) I felt that my knowledge of network engineering was lacking, so I specifically took some courses in network engineering to better understand how my code interacted with the network. You also brought up text encodings. Once at my job I did a full, rabbit-hole deep dive into text encodings and localization because we had frequent bugs related to localization, and I still find a lot of developers misunderstand a bunch of the important details in localization (e.g. the difference between an encoding and a unicode code point).

But sure, I don't need to understand every minutia of detail at the atomic detail. But I really like to know that I could, and in the meantime I have a enough of an understanding to build a complete conceptual model in my head.