Comment by oersted

16 hours ago

There's this dilemma where in theory there's a ton of demand for engineers that can do real LLM machine-learning, but in practice there are very few available positions and entrepreneurship opportunities.

The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.

Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.

With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.

In that regard, it's not too different from mobile telephony. Mobile phones drove the electronics industry 20 years ago, but there is limited demand for people who really know how to build a phone (ie. build the hardware and write all the signal processing from scratch), as there aren't that many companies that do phones at the lowest level. A few of the engineers got rich (eg. Viterbi), but most 'just' made a good living. Most people who got rich off phones didn't do it by knowing how phones work.

Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.

  • I was there at the start of the smartphone boom. I built a demo Android device that was capable of telephony/data, 3D rendering, etc. all before Google open-sourced the OS, for a SoC vendor that wasn't in Google's inner circle. Yet the industry was not interested in my junior profile during the subprime crisis.

  • This is super interesting because I moved from mobile telephony into ML and data science, and information theory and working with data in statistically correct way was what helped me! This was 10 years ago though.

  • Good points, but I think we can expect the AI space to be more tumultuous.

    What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way.

    What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.

    Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.

    • >* What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.*

      No.

      That’s what people like us on HN want. The people out in “Greater Userland” just want the black box to answer their questions. They could care less who is behind it. They don’t yet attach their black box to Amazon or Microsoft etc. And most won’t care enough to be inconvenienced even when they do make the connection. (As your competition argument implies.)

      Heck, a lot haven’t even made the connection between the black box that gives them answers and data centers. They think, “ ChatGPT good” and at the same time think “data centers bad”.

  • Same with telecom companies.

    We don't really have the demand for as many telecom companies as actually exist in the world. There's a reason we just have one Whatsapp and one Instagram, not three or four almost-but-not-quite clones in every single country that mostly differ in branding. The reason for the current situation has mostly to do with regulation and traditional, enterprise, "obviously every country needs a separate local branch, because that's what mcDonalds does" thinking. Technology has very little to do with it.

    This is why the telecom world now consist of equipment manufacturers, who do most of the hard tech stuff, and actual telecom companies, who operate the equipment, rig towers in their local country, and maybe write some glue code to integrate a core from vendor A, a billing system from vendor B and a CRM / corporate invoicing system from government-approved local vendor C.

    Banking also works similarly, though modern Neobanks / Fintechs and bank consolidation are slowly dissolving the concept of national bank branches.

  • Also, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.

Yes. It's like looking at the (Apollo) moon rocket launch and then suggesting teenagers should learn to build rockets in their garages for the coming space age.

It is viable as a toy project, but there are vanishingly few career opportunities.

  • It's like looking at the early internet and then suggesting teenagers should write browsers as their projects instead of webpages.

    • That may have worked at the time, but no companies are interested in learning projects today. If you didn’t do the reqs list for the last 5-10 years with the same title, forget about it. Because there are dozens of folks who have, lined up. No one is indulging career changers (and most fresh grads) for the time being.

      3 replies →

    • I mean... that's actually amazing advice. Not because they would grow up to create browser startups. But because they would grow up to create web startups that succeed because of very fundamental of how the web is rendered.

      Which is paulg's point really

      1 reply →

    • Early internet let me create the best personal web page in my city that I knew of with 2 weeks experience as a 12 year old. I imagine that same 12 year old could be more knowledgeable about LLMs than 99% of people in the same period.

      2 replies →

  • That seems like a perfectly good use of time for a kid in the 70’s! You’d learn a lot of engineering skills and demonstrate a tenaciousness that most people don’t have. I don’t think the goal is to predict what will be the important technology in 10 years. The goal is to challenge oneself with hard tasks and learn interesting stuff. A lot of the “AI” people today were compilers people yesterday

  • There are more career opportunities building rockets than designing rockets. Lots of welders, machinists, electrical engineers etc build rockets, and those skills transfer.

    The same is not so different for AI. A few people design novel AI, but there are a lot of people training AI (especially if you include fine tuning) and implementing AI, even as a hobby.

  • Building an LLM covers a lot of CS fundamentals and forces you to do lots of research in order to implement one, especially on more modest hardware.

    • So do a lot of things, that isn't what Paul is implying when he says "I would build a LLM if I was 17".

  • But he’s not saying that it’s a career opportunity.

    It seems to me like he’s saying that doing this thing would be 1) fun and 2) a great way to become employable in the future. I don’t believe he’s saying that this project would be some kind of job training exercise.

  • When I was 17 and still witnessing Apollo Moon launches, I wanted to build an AI that would handily outperform LLMs as we know them today.

    But that was way back in the early 1970's and all I had to work with was a mainframe.

    Well the mainframe itself wasn't bad, the real show-stopper was that I didn't own the computer outright, no strings attached, no debt, etc.

    >I'd probably try to make an LLM that I could use on some specific problem.

    I thought so too back then, still do so I guess this is one of those things that could stand the test of time. I always wanted to start with something a lot simpler than a Moon mission myself. At 17 I already had a significant breakthrough in the chem labs and it was from alternatives to a single processing step plus everything that descended from that, rather than trying to tackle a much more complex detailed multi-step synthesis. I was only 17 but I was not trying to be a slouch, I don't think pg was either at that age but his advice is not for just anybody. I couldn't have done it if I hadn't made major progress since being 16, and it really emphasized at the time how much maturity can make a difference. My imagination ran wild as I extrapolated :)

    In a reply from LeCun to pg:

    >>I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals. That last item is also what I would if I were 30, 40, 50, or 66 years old

    I see no reason to stop at 66 either ;)

    But I figured that people owning more computer power than I could ever afford were going to be doing something like this as soon as they could, without having to wait for something like an LLM to arrive before getting peoples' attention.

    It did seem like things were going to take longer than you expect, so it's pretty good to have a lifetime of concentrating on the specialized natural science domain expertise, focused now for 50 full years on how it would combine if AI ever got good enough.

    Both the natural science and the AI need to be a major cut above, I still see dramatic room for improvement in my own work. If I'm going to have to rely on "other peoples' AI" then that natural science component is going to have to pull a lot of weight to keep up with the kind of computers that only rich-as-hell high-rollers have access to.

I agree with mostly all of this, but personally I wrote a toy LLM almost 5 years ago and while it never saw much use outside of boring my wife with a shitty command line demo with glee it did help me understand how they worked and how to apply them, played a lot with JAX and pytorch, ended up building a ghetto version of MCP and an LLM-Pool to proxy requests to my baby local models and so I didn't struggle to see the evolution of openrouter and MCP agentic workflows. The same way i'm really glad when I was younger I built a bad webserver by myself, a really painful SQLx type database, etc etc etc - none of these things led me to developing for Nginx or Oracle nor will knowing JAX get me a job at an AI research lab, but I do have a lot of depth in understanding how the technology works so that the flavors on top of them are easy to digest and make more use of immediately, and I think the same can be said for engineers coming into the field - if it's a spooky LLM box you aren't going to be squeezing the same amount of juice as the guy that knows how they work inside and out so having at least the understanding of a _babys first LLM_ is going to get you miles ahead of people who don't.

For anyone who wants to dork around there is https://github.com/rasbt/LLMs-from-scratch which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.

  • Game cheating and reverse engineering MMO backends taught me a lot: databases, networking, securing a backend (and frontend), limitations of simpler languages when comparing them to more native options for building backends.

    • Agreed, I was very late to the game and was forced to learn VBA for excel sheets and that is how I finally broke into programming.

      When I was a pre-teen I stumbled upon CD-rom hacking guide to bypass disc requirements on games, I remember opening up the file and the screen being filled with HEX code. I was so overwhelmed I just closed it and never touched programming after that for 15 years. My life would have been totally different if I had embraced the unknown instead of retreating.

> The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.

And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like

> Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list} > 10+ years of experience in C/C++ > GPU programming experience required > 10 more ridiculous lines

No idea how a teenager self teaching himself LLMs is supposed to even get a shot...

  • It reminds me of that "stone soup" story.

    1. I can make turn a stone and water into a delicious soup

    "17 year olds, learn to build an LLM from scratch"

    2. This soup would be more delicious if we add a few carrots. Does anyone have carrots

    "Increase your chance of success by getting a Masters degree"

    3. How about potatoes?

    "And get a PHD"

    4. What about some salt?

    "And publish some peer reviewed articles in {journals_list}

    5. We should also add beef

    "Now work in the industry for 10 years"

    6. See, this soup is delicious, and I made it all with a stone

    "See, you're rich, and it's all because you learned LLMs as a 17 year old"

I have always been pro fundamentals. It caused me trouble early in my career with bosses that didn’t understand why I would spend time trying to understand how something worked at a low level if I was a high level user. But then knowing the fundamentals gave me an edge as a designer and developer by understanding capabilities and limitations of the tools I was using. For example understanding how indexes work internally in a relational database. So I see the value in this type of work, not to land a job as a LLM researcher, but as an informed user of the tool.

I’ve been wondering about this.

There are high school students competing in contests that cover parts of the (Math) theory behind AI. A lot of high school research programs are integrating AI with other things and complex mathematical models…

To me, this is bizarre as Calculus is barely taught in high schools (and likely poorly).

Don’t get me wrong, these kids certainly aren’t the usual lot.

Yet, I really wonder if they know the fundamentals. Do they even understand derivatives or just memorized the rule for polynomials? Can they even explain what a transistor is?

Normal curriculum takes 5 years to go from Algebra I to Calculus. Real Analysis, Linear Systems, etc. are fundamentals taught only in college…

Feels like too many are trying sprint before even learning to walk.

That's like saying the only way to do real engineering is with Google-scale Borg deployments. You can do quite a lot on very little hardware, r/StableDiffusion is a prime example.

  • You can do plenty of "real engineering" under normal conditions. But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large" for a reason. You can play around at small scale, but those lessons you learn will not be very relevant to the real problems in the market.

    Sure you can gradually climb the ladder by demonstrating your skills bit by bit and getting access to more resources. It has very good prospects if you do manage to push through. But it's a hard and risky path, and you will not be able to get any interesting results for the longest time.

    For a young middle-class student, it just doesn't make much sense. You can do much more impressive and impactful things with your time without getting into that black hole.

    I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. But they can't do much with it without capital or access.

    Good engineering has never been a bottleneck in this field, it's been all about having access to capital and taking smart but dangerous risks burning it on compute, without much idea of how long you need to keep burning for. There's still no end in sight, some are still managing to convince investors and keep burning, and we are seeing progress, but the business case is still unclear. If you want to get in that game, go ahead, but it's not something I would advice the average young engineer.

    • > I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex.

      I’m 40, and I don’t.I took that abstraction for granted and “left it to the big labs”. However I want to build my own LLM for learning purposes.

      On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have….

      The world needs pragmatic folks who work at a higher abstraction and make LLMs useful, AND also folks who think why not “this other way”? And build newer ways to do fundamental things.

      Given the usefulness of current LLMs, I would certainly encourage anybody to try and build their own LLMs, and see what they come up with…

      Heck if they build a rack full of old laptops and run something with it that could be done “better” with modern servers, I’d still appreciate the learning running things on those little machines bring.

      5 replies →

    • Agreed. It's hard to learn unless you have access to quite high end hardware, and even paying by the hour is expensive. There's a low ceiling on what you can learn without doing training runs.

      You can however learn everything you need to know to get on the career ladder as a software engineer on a regular home PC.

      1 reply →

    • > . But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large"

      The “large” qualifier dates back to pre-transformer language models, where even training a multi-million model was hard due to how poorly it scaled. GPT-2 was a large language model, despite being only 124 millions parameters.

      Due to how much high quality data is readily available, anyone can now train a sub-billion (L?)LM on commodity hardware.

      And I'm personally convinced that pretty much any enterprise use-case of an LLM (except coding) is better served by a fine-tuned small (<2B) model that is trained specifically on the task, rather than a generalist frontier model, so learning the engineering around fine-tuning is a key skill that companies will realize they need sooner than later.

      8 replies →

  • Interesting/capable diffusion models are much smaller than similarly interesting language models. But yes you could always scale things down to learn the fundamentals.

    • There are plenty of similarly tiny language models in the realm of tts too. Gatekeeping what’s interesting misses the forest for the trees.

I disagree with the premise.

Learning should not be done only as a direct path to getting paid.

Learn to create pattern matching and intuition to solve future problems.

When you are 17 it is a good time to understand how the world works so you can build on top of it in the future. If we assume most tech is going to have an LLM as part the stack, a solid basis in how LLMs work is likely to help you in future endeavors the same way a solid basis in how the web works helps you today.

Maybe a 17 year old should learn both. As a small anecdote when I was 17 I learned a lot about load balancers, failover, and building self-healing systems running small hosting company that had to be fault tolerant when I was attending high school. This wasn't at state of the art levels (e.g. I wasn't configuring gigabit routers or global CDNs -- but it was useful pattern matching for future problems)

I currently don't touch any of that tech, but I have working knowledge that still serves me today.

Think long term.

  • You haven't actually refuted their premise.

    • I am not the parent but I took it as them saying the premise was wrong to begin with, which I very much agree with. Learning should not be primarily directed by job availability.

A lot of startup companies are not training frontier models but help solve and optimize pain points of LLMs: cyber security, token usage, harnesses etc. These jobs don't require a PHD in machine learning but it does help if you understand LLMs at a deeper level.

  • A lot of startups succeed by lying to themselves about the quality of their solutions and focusing on selling a marketing story.

    • True, I'm tempted to put Antrhopic under the same category of lying startups so it enforces my point that understanding LLMs deeply could be useful

But you can train a small LLM with a gaming graphics card -- I managed one on a GTX 1660. I don't think pg is suggesting that you try to chase the frontier. It's more like building your own OS in the 80s, or web server in the 90s -- sure, you'll never match the commercial offerings or the big OS projects, but building something from scratch within the limits of the hardware you can afford is amazing educationally.

  • "I'd build the foundation of knowledge to base a startup on later"

    what kinds of startups ?

    • In my experience, having a solid understanding of the next level down in the stack -- the foundation you're building your startup on -- is really helpful. We built a PaaS, and knowing enough about Linux internals to be able to work out what would be easy and what would be hard meant that we could focus our efforts on high bang-for-buck features.

      So I'd say that understanding LLMs to the level that you get to by training your own baby one would be a solid foundation for pretty much anything built on top of the "real" ones.

Y'all are missing the point: It's probably less than 100 hours to learn the foundations of one of humanity's most-current breakthroughs. It's a disservice to any young hacker to not learn it. Here's your curriculum. Watch these:

-- 3Blue 1Brown's Neural Network Series: https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700...

-- Karpathy on LLMs: https://www.youtube.com/watch?v=7xTGNNLPyMI

-- Stanford CS336: https://www.youtube.com/watch?v=JuoVZkPBiKk

Then do this hands-on:

-- Karpathy's zero-to-hero: https://karpathy.ai/zero-to-hero.html

when you think about all of the advancements since Attention / GPT a lot of it has been somewhat more obvious than in other fields, as is typical with the massive flood of innovation that follows a big breakthrough.

Paul likely assumes there will be a sequence of additional papers with the same impact as Attention is all you need, which will spawn a lot of opportunity for a larger group of experts who are conversant enough to advance the field even if they do not themselves create such a major innovation. Not only is this deeply exciting, it is also highly meritocratic as there is still scarcity of the kind of intellect and creativity necessary to swim there.

Machine intelligence might soon surpass it, though, and Deepseek is 100% Chinese mainland educated. Paul's description of building an LLM from scratch is meant as a vague starting point for being an innovator of the highest value aspect of modern AI innovation, not as a specific prescription.

AI is the subtrate the future runs on.

And so I think the idea is more to understand tomorrow ... from first principles.

In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me.

I haven't written assembly in years.

But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.

  • > AI is the subtrate the future runs on.

    We don't know this. So many people are simply claiming this confidently, and a lot of them are betting their careers on it, but nobody has a crystal ball. Whenever someone tells you confidently, and without any doubt or qualifications, that something "is the future," be skeptical.

    I remember when the Segway was definitely going to change urban planning worldwide.

    • Segways failed because they were a solution looking for a problem. You can't say that about current-gen LLMs, much less AI in general.

  • > AI is the subtrate the future runs on.

    Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.

    So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.

    But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.

    And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.

    We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.

    • > But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans ... without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.

      At 75% replacement of a worker we would already have huge job losses as each individual would be doing what several before did.

      The only alleviation would be the creation of new equivalently paid jobs, which is no better than a hypothesis right now.

    • I keep saying AI is going to prove more impactful than cloud computing but less impactful than the sewing machine and a bunch of people get mad at me.

  • I don't really agree, I think the future should run on humans.

    AI is a great solution in niche areas but generally doesn't make much money. All the large companies are in the negative.

    The steam engine was less of a bubble and was much more revolutionary and had a greater impact.

A single 3090 will train qwen 0.8B just fine. While it’s not a very capable model any training technique you would want to master can be used to make real progress. And all the skills you need to learn how to do this can be learned watching Andrej Karpathy’s zero to hero series (shame he quit educational content and went to anthropic)

This is a wildly incorrect and myopic view on the world.

Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.

There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.

  • I find the fine tune approach more interesting than straight to RAG and MCP.

    End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes.

    Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient.

    I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.

  • Both of you are right. There is demand for tailored (fine-tuned) models; almost every enterprise would theoretically benefit from them.

    But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not!

    Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.

We finetune LLMs. Small ones like Gemma 4 for semantic tasks.

There are plenty of areas were we need people to do this for insurances, banks etc.

AI/ML exists on many levels.

Yes and no. I believe the point he is making is simply that there is no substitute for fundamentals and first-principles thinking.

We had scores of students study how microprocessors work and compilers work over decades, yet we have 3 or 4 major processor companies and a handful of programming languages. Yet, what they learned was probably crucial in their development as engineers.

We are also so early right now that even 2-3 years from now who knows how many LLMs and model firms survive (esp. given the "snake eating its tail" venture/investor funding situation)

  • This is kind of different though isn't it? Doing an assembly or compiler class has pretty clear benefits in this regard.

    But LLMs are tools. Does a great engineer need to know how vscode works? Might be helpful to understand how extensions work, LSPs, and project configurations.

    Usually when working with any tools, you need to understand how to get the most out of your tool for your needs and that's about it. Core fundamentals about how software and hardware works in general seems like it would be MUCH more useful than LLM core knowledge.

I think his point is to do this to understand deeply what they can do, what they can’t do, and what they can almost do. And then find the highest value ‘almost’ use case and push there. Which doesn’t necessarily mean improve the llm, could be applying it in just the right way for the use case. Of course, the bitter lesson makes this hard and risky. But no more risky than investing your time in learning anything else these days.

"Necessity is the mother of invention" - limited hardware has always forced people to find cleverer ways of doing more with less. Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).

  • This is roughly how GPUs for neural networks got started: after Andrew Ng left Google Brain, he no longer had access to a 10,000-CPU cluster used to train the original DistBelief system. But his Stanford students could buy a GPU...

  • > Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).

    I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms.

    And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.

I bet this will get less true over time though as the rate of change slows down, allowing specialized models/training for specific use cases that aren't TAM heavy enough for the big labs to go after them. It's just now any general model is the best thing to use for everything and you're wasting money to build something on what will certainly be obsolete by the time you can get it to market

I don't think pg is giving advice on what will lead most directly to a job, but rather what is the best learning for a 17yo.

A 17yo who trains their own LLM will have a much richer understanding of what AI is, how it works, what its potential capabilities and pitfalls are, versus someone who spends the same time doing something else.

You can train and run small models on an old gpu. That’s what I’m doing now at, well, much older than 17. Does it produce a useful model? No. Not even remotely.

However, I do learn stuff about models that takes it from “magic” to “useful tool I understand the limitations of.”

Do I do it for that reason? No not really, I’ve never had luck learning something because it would be good for my career. I do it because at my core I’m a bored teenager who wants to make the computer do cool shit.

It is not a skill that you will use in your day to day life, but I think it is part of the fundamentals now. Sure, LLMs are in a bubble, just like the web during the dotcom bubble, but web didn't disappear, and I don't expect LLMs to, even after the bubble bursts.

I didn't write a LLM from scratch but it is on my "wishlist" so to speak. From what it seems, a GPT-1 class LLM can be done from scratch in a few days and tens of dollars of cloud compute or a high-end gaming GPU.

It is an exercise not unlike building a compiler, a school classic. You are unlikely to ever work on a compiler, but at least, now, you know your tools a little better. It is not about becoming an expert, that takes years, it is about knowing what you are doing.

If you intend to make software engineering your career, you will want more than surface knowledge. And that part is entirely on you, or on your school if you are a student. Companies will not pay for you to learn the fundamentals, they want short term returns, because you may leave at any time. But you as a software engineer may have 40+ years left, so it is worth thinking long term. Claude code may become obsolete a few years, but linear algebra is not going anywhere.

It is on my list to build a toy LLM from scratch.

Not that expect to make it big as a LLM researcher but building something from scratch gives a much deeper understanding than what you can get from simply using something.

Much in the same way as implementing and designing your own programming language makes you a much better programmer.

Of course, because it's not the LLM that's special but the training data. Nowadays, your favourite AI service to generate code for an LLM whenever you ask for it.

  • I don’t think that’s correct, the data is not that special either, and getting a similar dataset is significantly easier than getting the compute capacity to use it, even if they are both relatively hard.

    Probably this also is too cynical and simplistic, but: really what’s special is the ability to get this kind of capital, with the freedom to burn it on mad moonshots, with long enough leeway to actually get to see a few of the moonshots come true.

    No wonder that the head of YC made this happen, this is exactly what they are world-leading at.

What about other machine learning related skills? Does this wave of LLM mean less need for that kind of work?

I would think not, but when I started to look into OCR options recently - assuming that obviously a dedicated tool would do a better job than an LLM - I was wrong (apparently).

+1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.

The only jobs that he found he was highly over qualified or paid very little.

In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.

  • > +1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.

    Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.

    • How've you gone about getting into this btw? I have extensive experience in infra and pipeline rollout but have struggled to find freelance clients for this kind of thing. Would be great to tie it into ML as a learning opportunity there

      2 replies →

    • This only proves the original point which is that there is not much demand for actual machine learning expertise because that is only carried out in a small number of places and what demands there is is for the more basic software carpentry like infrastructure and operations rather than the actual technology and Engineering side of things

      3 replies →

  • I mean it makes sense right? Anthropic for instance has like, a couple hundred staff in London with plans to expand to somewhere just shy of a thousand. There are far far more ML/Maths/CS/Stats PhD's than there are openings. Especially in London there is no shortage of suitable candidates given Cambridge/Oxford/Imperial/UCL are surrounding it. 2% of the UK population has a PhD alone...

I think a lot of this is based on preconceptions. A lot of apps were made with Electron, because it was common wisdom that native is 'too hard'.

Now with LLMs, people write native apps in Rust, and I'd like to think some of them found that there isn't such a huge jump in difficulty they assumed there would be.

  • It never was native "too hard" it was always "too expensive".

    That's the same case finding companies that will actually pay for hand made LLM instead of using something from big providers will be hard because most companies won't be able to afford it.

    Yeah if you find a company that will do that stuff directly, good for you, but you will have to be very lucky and you will have to compete with other people who followed PG advice.

    So I would rather learn all there is about properly using LLMs and integrating them with existing systems, that will most likely by useful for 90% of companies out there.

    Building business niche harnesses is in my opinion much better direction. Knowing what will work best in specific cases is it FTS or vector search, optimisation of usage, getting best results while using cheaper models, knowing how to use tools to run models on the servers, and all the tooling around that like various MCP or just tooling that will be provided to models.

    That is what I am currently busy with and I already have customers for that knowledge.

  • Oh yes I agree, LLMs are not that complex in principle, most engineers could build a toy version completely from scratch without too much difficulty.

    But that’s the tip of the iceberg. If you have any ambitions of doing this professionally, it quickly becomes clear that all it’s all about knowing how to deal with problems that are only present at massive scale, when an LLM is actually L and becomes AI.

    The mundane details about how to build a tiny autocomplete model and the maths behind it you can learn in a couple weeks easily. It’s not black magic, there are much harder areas of computers science.

This is a silly take. You can learn to build an LLM, there are great resources to do so (there are books about building them from scratch), you can use older model GPUs or rent them by the hour. The value of understanding them is really high for anyone building any application that uses an LLM at any point.

It similar to understanding how a very basic CPU works. Just because I'm not going to work at intel or nvidia or whatever optimizing the hell out of a chip, it doesn't mean I just throw my hands up and think "magic" - the basic architecture isn't that difficult, and the value of knowing it is astronomical for anyone writing software.

FWIW, at least 20 Y Combinator startups have published ML research recently at ICLR, NeurIPS, ICML, and so on.

I think a lot of people assume that only the big AI labs can do cutting edge research, but there's a strong argument you can do it as part of little tech as well.

> With all due respect, paulg is in somewhat of a bubble

I feel like the "ALWAYS HAS BEEN" meme is apropos here.

lol unnecessary and inefficient…

I just finished fine tuning Gemma e2b for local code completion on my local machine.

This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments

  • That’s great but the industry has moved on to code generation and editing, not code completion.

    • 1. I very much still write code myself without an LLM when I need top quality

      2. That's why I have an agentic agent as well installed, Qwen 27B, outrageously good, better than sonnet 2 years ago. And it's mine, I can give it confidential info to work with since I own the whole chain. See where I'm going with this?

  • Can you tell me what specs your machine has? There is a difference between a few hours and a few days

    • A MBP m5 in this case, but the fine tune would have cost me like 10 bucks on RunPod, that's what I did with my previous setup.

The other thing to add as well is that the research teams who do the actual research work are relatively small and very specifically qualified which naturally keeps the barrier to entry high.

I think companies of all sizes will want their own models, or at least customised ones, for their own specific use cases or competition and security issues.

1. Both training and optimisation will get significantly cheaper and easier quickly.

2. Politics will probably get even more insane before a potential reprieve on the 20th of Jan 2029.

3. The big AI firms will become part of the surveillance capitalism network, if they're not already.

So I think for self-protection a lot of companies will be looking near to medium term AI independence.

  • The argument is sound, but the maths don't math for now, and it's unclear when/if they will.

    For the time being, unless you truly have millions, the outcome from training will be very net negative, while focusing on building on top of existing AI will yield amazing things if you apply the same talent and effort.

    When it does get cheaper, then it will be easier to acquire the skills and experience too, and the struggle you went through by trying to do it now will be somewhat wasted.

    Besides, I am well versed in this field, and it is not rocket science. There are plenty of software engineering domains that are a lot more challenging, like high-end graphics, large-scale data engineering or kernel programming. People will learn to train LLMs when people want them to.

  • Right, just like companies don't use SAAS.

    In reality, enterprises are happy to offload even risky tasks to others as long as they get some contractual guarantees about their data. Would they like more choice in who to buy from? Yes, but not enough to in-house such a specific discipline.

  • The cost of training a model from scratch is going to be cost prohibitive for the vast majority of companies (even if renting the hardware needed for the 1-2 month training time). It's an interesting learning exercise, and some of the things learned can be applied to other parts of the process. There's also the issue of needing a huge amount of data needed to get decent weights.

    Fine-tuning a model or LoRA based on the companies data set is more feasible but you're likely going to need several runs as you test/try out different base models, parameters, etc. This is why there are a lot of fine-tuned models on huggingface based on base or instruction-trained models from the larger AI companies that have released open weight models (Microsoft, Google, IBM, Mistral, DeepSeek, Qwen, etc.).

    Training is limited on memory first (storing training data and weights) and computation second. Realistically you need to own or rent 2-8 H100/B100 devices or Google's TPUs.

    The majority of workflows for a company providing AI capabilities are likely best solved by tailoring a system prompt for the chosen model, evaluating the prompt and model with tools like promptfoo, and then running it on a compute cloud provider (including AWS Bedrock). If the company is big/financially well off enough they could look at buying the hardware needed to run it on their own servers.

    For other uses like agentic software development you'd need to spin up a suitable model on a compute cloud provider (or local hardware if the model is small enough) and then tell your IDE/editor to use that model. You would need some way of benchmarking and evaluating the models to see if they are capable of doing the tasks you need. -- There have been some tests done by people on YouTube that suggests that Qwen 3.8 27B is a decent model, but your needs may vary.

    • Even for most organizations, testing AI systems is too cost prohibitive, so they YOLO in production, including public facing systems.

  • Most companies that build physical goods don't care for one second about their IT department other than how much money they can save per month, starting by outsourcing whole of it, thus they have little use for internal LLMs.

    • And it's across the industry, thinking banks, private banks, insurance, pharamcy etc don't outsource their IT, including development... I believe US outsource even more than Europe on this matter.

> With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.

Paul, I think, is talking about achieving outsized outcomes in relatively shorter timeframes (as the timing is just right to be investing in learning this tech) for high agency folks who can also afford the ordeal in wanting to maximize for impact & ambition. Of course, there's real risk one may get no where, but even in failure, given you were building the LLM yourself, you might end up with other adjacent, high reward opportunities.

> The reality is that an incredibly small minority of companies in the world do any real training or optimisation.

At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.

Did you read his comments on this? It's not to actually do LLM research stuff, it's to trigger and unlock ideas.

I think his point is, if LLMs are the future (like computing is the future in the 90s), you should be an LLM-expert (equivalent of becoming a software developer).

I can see the point. It's unlikely that a 2.4T LLM will be integrated into, say, a pesticide drone. You'll still need some kind of LLMs to achieve maneuvers that "normal" programmings can't achieve.

But what if everything basically turn into that? Essentially, instead of build me a web app to solve X and do Y, build me an LLM to serve X and do Y. (unless the current LLMs are able to do it end-to-end but then they can hardly write coherent software/personal opinion).

This is like saying, “teens shouldn’t learn how to make their own game engine because no one is hiring for that.”

You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.

  • > You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.

    Writing your own game engine makes you realize that the Unity engine is not really that well written....

    • Writing your own game engine makes you realise that no game engine is written well (when it is written to ship a game).

  • Except knowing how LLMs work don’t actually provide much understanding for using them. People don’t use LLMs the way we’ve built on most other tons or platforms. It’s more learning Unity hoping to be a better gamer.

That's not true, because everyone, everyone, everyone seems to want to do training. Which results in a 50 person company training, say, a voice model that then fails, because it's just not good enough.

In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human)

... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels.

e.g. https://valueaddvc.com/blog/inference-chips-vs-training-chip... (I get it, not great first source, but st

  • The big question is whether companies hold enough proprietary data to do useful things that for e.g. Anthropic, etc. can't easily replicate.

    For some very niche cases I think this is probably the case but for the vast majority, the company's data isn't as useful as they think it is or anywhere near the size needed.

  • Everyone says they want to do training, because it's sexy and an easy way to justify raising mad funding rounds. Some manage, most don't.

    I don't know where you are located, but in EU, in China, and yes even in Silicon Valley, the vast majority of companies do not do any real AI engineering. There's nothing wrong with it, it's just not a smart path for most purposes. You can do amazing things without training, and if you try to train, you cannot get anything amazing unless you burn millions.

    Very few people can afford to play the long game and cross that dessert. And, sure, you will not get far without good engineering, but good engineering is definitely not sufficient and is not the primary bottleneck.

    • It took a long time to cross that desert, and no sane company would want to get stuck in a desert, unless it's specifically an R+D "desert crossing" company.