I were 17, I'd learn how to build LLMs from scratch

20 hours ago (twitter.com)

https://xcancel.com/paulg/status/2091544343589060625

There's this dilemma where in theory there's a ton of demand for engineers that can do real LLM machine-learning, but in practice there are very few available positions and entrepreneurship opportunities.

The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.

Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.

With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.

  • In that regard, it's not too different from mobile telephony. Mobile phones drove the electronics industry 20 years ago, but there is limited demand for people who really know how to build a phone (ie. build the hardware and write all the signal processing from scratch), as there aren't that many companies that do phones at the lowest level. A few of the engineers got rich (eg. Viterbi), but most 'just' made a good living. Most people who got rich off phones didn't do it by knowing how phones work.

    Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.

    • I was there at the start of the smartphone boom. I built a demo Android device that was capable of telephony/data, 3D rendering, etc. all before Google open-sourced the OS, for a SoC vendor that wasn't in Google's inner circle. Yet the industry was not interested in my junior profile during the subprime crisis.

      4 replies →

    • Good points, but I think we can expect the AI space to be more tumultuous.

      What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way.

      What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.

      Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.

    • This is super interesting because I moved from mobile telephony into ML and data science, and information theory and working with data in statistically correct way was what helped me! This was 10 years ago though.

    • Also, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.

  • Yes. It's like looking at the (Apollo) moon rocket launch and then suggesting teenagers should learn to build rockets in their garages for the coming space age.

    It is viable as a toy project, but there are vanishingly few career opportunities.

    • That seems like a perfectly good use of time for a kid in the 70’s! You’d learn a lot of engineering skills and demonstrate a tenaciousness that most people don’t have. I don’t think the goal is to predict what will be the important technology in 10 years. The goal is to challenge oneself with hard tasks and learn interesting stuff. A lot of the “AI” people today were compilers people yesterday

    • There are more career opportunities building rockets than designing rockets. Lots of welders, machinists, electrical engineers etc build rockets, and those skills transfer.

      The same is not so different for AI. A few people design novel AI, but there are a lot of people training AI (especially if you include fine tuning) and implementing AI, even as a hobby.

    • When I was 17 and still witnessing Apollo Moon launches, I wanted to build an AI that would handily outperform LLMs as we know them today.

      But that was way back in the early 1970's and all I had to work with was a mainframe.

      Well the mainframe itself wasn't bad, the real show-stopper was that I didn't own the computer outright, no strings attached, no debt, etc.

      >I'd probably try to make an LLM that I could use on some specific problem.

      I thought so too back then, still do so I guess this is one of those things that could stand the test of time. I always wanted to start with something a lot simpler than a Moon mission myself. At 17 I already had a significant breakthrough in the chem labs and it was from alternatives to a single processing step plus everything that descended from that, rather than trying to tackle a much more complex detailed multi-step synthesis. I was only 17 but I was not trying to be a slouch, I don't think pg was either at that age but his advice is not for just anybody. I couldn't have done it if I hadn't made major progress since being 16, and it really emphasized at the time how much maturity can make a difference. My imagination ran wild as I extrapolated :)

      In a reply from LeCun to pg:

      >>I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals. That last item is also what I would if I were 30, 40, 50, or 66 years old

      I see no reason to stop at 66 either ;)

      But I figured that people owning more computer power than I could ever afford were going to be doing something like this as soon as they could, without having to wait for something like an LLM to arrive before getting peoples' attention.

      It did seem like things were going to take longer than you expect, so it's pretty good to have a lifetime of concentrating on the specialized natural science domain expertise, focused now for 50 full years on how it would combine if AI ever got good enough.

      Both the natural science and the AI need to be a major cut above, I still see dramatic room for improvement in my own work. If I'm going to have to rely on "other peoples' AI" then that natural science component is going to have to pull a lot of weight to keep up with the kind of computers that only rich-as-hell high-rollers have access to.

    • But he’s not saying that it’s a career opportunity.

      It seems to me like he’s saying that doing this thing would be 1) fun and 2) a great way to become employable in the future. I don’t believe he’s saying that this project would be some kind of job training exercise.

      2 replies →

  • I agree with mostly all of this, but personally I wrote a toy LLM almost 5 years ago and while it never saw much use outside of boring my wife with a shitty command line demo with glee it did help me understand how they worked and how to apply them, played a lot with JAX and pytorch, ended up building a ghetto version of MCP and an LLM-Pool to proxy requests to my baby local models and so I didn't struggle to see the evolution of openrouter and MCP agentic workflows. The same way i'm really glad when I was younger I built a bad webserver by myself, a really painful SQLx type database, etc etc etc - none of these things led me to developing for Nginx or Oracle nor will knowing JAX get me a job at an AI research lab, but I do have a lot of depth in understanding how the technology works so that the flavors on top of them are easy to digest and make more use of immediately, and I think the same can be said for engineers coming into the field - if it's a spooky LLM box you aren't going to be squeezing the same amount of juice as the guy that knows how they work inside and out so having at least the understanding of a _babys first LLM_ is going to get you miles ahead of people who don't.

    For anyone who wants to dork around there is https://github.com/rasbt/LLMs-from-scratch which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.

    • Game cheating and reverse engineering MMO backends taught me a lot: databases, networking, securing a backend (and frontend), limitations of simpler languages when comparing them to more native options for building backends.

      3 replies →

  • > The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.

    And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like

    > Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list} > 10+ years of experience in C/C++ > GPU programming experience required > 10 more ridiculous lines

    No idea how a teenager self teaching himself LLMs is supposed to even get a shot...

  • I disagree with the premise.

    Learning should not be done only as a direct path to getting paid.

    Learn to create pattern matching and intuition to solve future problems.

    When you are 17 it is a good time to understand how the world works so you can build on top of it in the future. If we assume most tech is going to have an LLM as part the stack, a solid basis in how LLMs work is likely to help you in future endeavors the same way a solid basis in how the web works helps you today.

    Maybe a 17 year old should learn both. As a small anecdote when I was 17 I learned a lot about load balancers, failover, and building self-healing systems running small hosting company that had to be fault tolerant when I was attending high school. This wasn't at state of the art levels (e.g. I wasn't configuring gigabit routers or global CDNs -- but it was useful pattern matching for future problems)

    I currently don't touch any of that tech, but I have working knowledge that still serves me today.

    Think long term.

  • I have always been pro fundamentals. It caused me trouble early in my career with bosses that didn’t understand why I would spend time trying to understand how something worked at a low level if I was a high level user. But then knowing the fundamentals gave me an edge as a designer and developer by understanding capabilities and limitations of the tools I was using. For example understanding how indexes work internally in a relational database. So I see the value in this type of work, not to land a job as a LLM researcher, but as an informed user of the tool.

  • I’ve been wondering about this.

    There are high school students competing in contests that cover parts of the (Math) theory behind AI. A lot of high school research programs are integrating AI with other things and complex mathematical models…

    To me, this is bizarre as Calculus is barely taught in high schools (and likely poorly).

    Don’t get me wrong, these kids certainly aren’t the usual lot.

    Yet, I really wonder if they know the fundamentals. Do they even understand derivatives or just memorized the rule for polynomials? Can they even explain what a transistor is?

    Normal curriculum takes 5 years to go from Algebra I to Calculus. Real Analysis, Linear Systems, etc. are fundamentals taught only in college…

    Feels like too many are trying sprint before even learning to walk.

  • Of course, because it's not the LLM that's special but the training data. Nowadays, your favourite AI service to generate code for an LLM whenever you ask for it.

    • I don’t think that’s correct, the data is not that special either, and getting a similar dataset is significantly easier than getting the compute capacity to use it, even if they are both relatively hard.

      Probably this also is too cynical and simplistic, but: really what’s special is the ability to get this kind of capital, with the freedom to burn it on mad moonshots, with long enough leeway to actually get to see a few of the moonshots come true.

      No wonder that the head of YC made this happen, this is exactly what they are world-leading at.

  • when you think about all of the advancements since Attention / GPT a lot of it has been somewhat more obvious than in other fields, as is typical with the massive flood of innovation that follows a big breakthrough.

    Paul likely assumes there will be a sequence of additional papers with the same impact as Attention is all you need, which will spawn a lot of opportunity for a larger group of experts who are conversant enough to advance the field even if they do not themselves create such a major innovation. Not only is this deeply exciting, it is also highly meritocratic as there is still scarcity of the kind of intellect and creativity necessary to swim there.

    Machine intelligence might soon surpass it, though, and Deepseek is 100% Chinese mainland educated. Paul's description of building an LLM from scratch is meant as a vague starting point for being an innovator of the highest value aspect of modern AI innovation, not as a specific prescription.

  • It is not a skill that you will use in your day to day life, but I think it is part of the fundamentals now. Sure, LLMs are in a bubble, just like the web during the dotcom bubble, but web didn't disappear, and I don't expect LLMs to, even after the bubble bursts.

    I didn't write a LLM from scratch but it is on my "wishlist" so to speak. From what it seems, a GPT-1 class LLM can be done from scratch in a few days and tens of dollars of cloud compute or a high-end gaming GPU.

    It is an exercise not unlike building a compiler, a school classic. You are unlikely to ever work on a compiler, but at least, now, you know your tools a little better. It is not about becoming an expert, that takes years, it is about knowing what you are doing.

    If you intend to make software engineering your career, you will want more than surface knowledge. And that part is entirely on you, or on your school if you are a student. Companies will not pay for you to learn the fundamentals, they want short term returns, because you may leave at any time. But you as a software engineer may have 40+ years left, so it is worth thinking long term. Claude code may become obsolete a few years, but linear algebra is not going anywhere.

  • AI is the subtrate the future runs on.

    And so I think the idea is more to understand tomorrow ... from first principles.

    In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me.

    I haven't written assembly in years.

    But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.

    • > AI is the subtrate the future runs on.

      We don't know this. So many people are simply claiming this confidently, and a lot of them are betting their careers on it, but nobody has a crystal ball. Whenever someone tells you confidently, and without any doubt or qualifications, that something "is the future," be skeptical.

      I remember when the Segway was definitely going to change urban planning worldwide.

    • I don't really agree, I think the future should run on humans.

      AI is a great solution in niche areas but generally doesn't make much money. All the large companies are in the negative.

      The steam engine was less of a bubble and was much more revolutionary and had a greater impact.

    • > AI is the subtrate the future runs on.

      Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.

      So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.

      But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.

      And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.

      We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.

      2 replies →

  • But you can train a small LLM with a gaming graphics card -- I managed one on a GTX 1660. I don't think pg is suggesting that you try to chase the frontier. It's more like building your own OS in the 80s, or web server in the 90s -- sure, you'll never match the commercial offerings or the big OS projects, but building something from scratch within the limits of the hardware you can afford is amazing educationally.

  • That's like saying the only way to do real engineering is with Google-scale Borg deployments. You can do quite a lot on very little hardware, r/StableDiffusion is a prime example.

    • You can do plenty of "real engineering" under normal conditions. But specifically when it comes to LLM engineering, no there's really not much you can do, they are called "large" for a reason. You can play around at small scale, but those lessons you learn will not be very relevant to the real problems in the market.

      Sure you can gradually climb the ladder by demonstrating your skills bit by bit and getting access to more resources. It has very good prospects if you do manage to push through. But it's a hard and risky path, and you will not be able to get any interesting results for the longest time.

      For a young middle-class student, it just doesn't make much sense. You can do much more impressive and impactful things with your time without getting into that black hole.

      I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. But they can't do much with it without capital or access.

      Good engineering has never been a bottleneck in this field, it's been all about having access to capital and taking smart but dangerous risks burning it on compute, without much idea of how long you need to keep burning for. There's still no end in sight, some are still managing to convince investors and keep burning, and we are seeing progress, but the business case is still unclear. If you want to get in that game, go ahead, but it's not something I would advice the average young engineer.

      16 replies →

    • Interesting/capable diffusion models are much smaller than similarly interesting language models. But yes you could always scale things down to learn the fundamentals.

      1 reply →

  • A lot of startup companies are not training frontier models but help solve and optimize pain points of LLMs: cyber security, token usage, harnesses etc. These jobs don't require a PHD in machine learning but it does help if you understand LLMs at a deeper level.

  • This is a wildly incorrect and myopic view on the world.

    Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.

    There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.

    • I find the fine tune approach more interesting than straight to RAG and MCP.

      End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes.

      Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient.

      I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.

    • Both of you are right. There is demand for tailored (fine-tuned) models; almost every enterprise would theoretically benefit from them.

      But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not!

      Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.

  • I don't think pg is giving advice on what will lead most directly to a job, but rather what is the best learning for a 17yo.

    A 17yo who trains their own LLM will have a much richer understanding of what AI is, how it works, what its potential capabilities and pitfalls are, versus someone who spends the same time doing something else.

  • I think his point is to do this to understand deeply what they can do, what they can’t do, and what they can almost do. And then find the highest value ‘almost’ use case and push there. Which doesn’t necessarily mean improve the llm, could be applying it in just the right way for the use case. Of course, the bitter lesson makes this hard and risky. But no more risky than investing your time in learning anything else these days.

  • I bet this will get less true over time though as the rate of change slows down, allowing specialized models/training for specific use cases that aren't TAM heavy enough for the big labs to go after them. It's just now any general model is the best thing to use for everything and you're wasting money to build something on what will certainly be obsolete by the time you can get it to market

  • Yes and no. I believe the point he is making is simply that there is no substitute for fundamentals and first-principles thinking.

    We had scores of students study how microprocessors work and compilers work over decades, yet we have 3 or 4 major processor companies and a handful of programming languages. Yet, what they learned was probably crucial in their development as engineers.

    We are also so early right now that even 2-3 years from now who knows how many LLMs and model firms survive (esp. given the "snake eating its tail" venture/investor funding situation)

    • This is kind of different though isn't it? Doing an assembly or compiler class has pretty clear benefits in this regard.

      But LLMs are tools. Does a great engineer need to know how vscode works? Might be helpful to understand how extensions work, LSPs, and project configurations.

      Usually when working with any tools, you need to understand how to get the most out of your tool for your needs and that's about it. Core fundamentals about how software and hardware works in general seems like it would be MUCH more useful than LLM core knowledge.

  • It is on my list to build a toy LLM from scratch.

    Not that expect to make it big as a LLM researcher but building something from scratch gives a much deeper understanding than what you can get from simply using something.

    Much in the same way as implementing and designing your own programming language makes you a much better programmer.

  • "Necessity is the mother of invention" - limited hardware has always forced people to find cleverer ways of doing more with less. Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).

    • This is roughly how GPUs for neural networks got started: after Andrew Ng left Google Brain, he no longer had access to a 10,000-CPU cluster used to train the original DistBelief system. But his Stanford students could buy a GPU...

    • > Current models are clearly nowhere near the efficiency limit (the brain does vastly more with far less power).

      I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms.

      And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.

  • What about other machine learning related skills? Does this wave of LLM mean less need for that kind of work?

    I would think not, but when I started to look into OCR options recently - assuming that obviously a dedicated tool would do a better job than an LLM - I was wrong (apparently).

  • You can train and run small models on an old gpu. That’s what I’m doing now at, well, much older than 17. Does it produce a useful model? No. Not even remotely.

    However, I do learn stuff about models that takes it from “magic” to “useful tool I understand the limitations of.”

    Do I do it for that reason? No not really, I’ve never had luck learning something because it would be good for my career. I do it because at my core I’m a bored teenager who wants to make the computer do cool shit.

  • I think a lot of this is based on preconceptions. A lot of apps were made with Electron, because it was common wisdom that native is 'too hard'.

    Now with LLMs, people write native apps in Rust, and I'd like to think some of them found that there isn't such a huge jump in difficulty they assumed there would be.

    • It never was native "too hard" it was always "too expensive".

      That's the same case finding companies that will actually pay for hand made LLM instead of using something from big providers will be hard because most companies won't be able to afford it.

      Yeah if you find a company that will do that stuff directly, good for you, but you will have to be very lucky and you will have to compete with other people who followed PG advice.

      So I would rather learn all there is about properly using LLMs and integrating them with existing systems, that will most likely by useful for 90% of companies out there.

      Building business niche harnesses is in my opinion much better direction. Knowing what will work best in specific cases is it FTS or vector search, optimisation of usage, getting best results while using cheaper models, knowing how to use tools to run models on the servers, and all the tooling around that like various MCP or just tooling that will be provided to models.

      That is what I am currently busy with and I already have customers for that knowledge.

    • Oh yes I agree, LLMs are not that complex in principle, most engineers could build a toy version completely from scratch without too much difficulty.

      But that’s the tip of the iceberg. If you have any ambitions of doing this professionally, it quickly becomes clear that all it’s all about knowing how to deal with problems that are only present at massive scale, when an LLM is actually L and becomes AI.

      The mundane details about how to build a tiny autocomplete model and the maths behind it you can learn in a couple weeks easily. It’s not black magic, there are much harder areas of computers science.

  • FWIW, at least 20 Y Combinator startups have published ML research recently at ICLR, NeurIPS, ICML, and so on.

    I think a lot of people assume that only the big AI labs can do cutting edge research, but there's a strong argument you can do it as part of little tech as well.

  • We finetune LLMs. Small ones like Gemma 4 for semantic tasks.

    There are plenty of areas were we need people to do this for insurances, banks etc.

    AI/ML exists on many levels.

  • The other thing to add as well is that the research teams who do the actual research work are relatively small and very specifically qualified which naturally keeps the barrier to entry high.

  • +1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.

    The only jobs that he found he was highly over qualified or paid very little.

    In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.

    • > +1, I have a friend, math PhD that's been working on ML research 5+ years in London yet he has not been able to find any position.

      Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.

      8 replies →

    • As always and everywhere, it's who you know (and who knows you) that matters.

    • I mean it makes sense right? Anthropic for instance has like, a couple hundred staff in London with plans to expand to somewhere just shy of a thousand. There are far far more ML/Maths/CS/Stats PhD's than there are openings. Especially in London there is no shortage of suitable candidates given Cambridge/Oxford/Imperial/UCL are surrounding it. 2% of the UK population has a PhD alone...

      2 replies →

  • lol unnecessary and inefficient…

    I just finished fine tuning Gemma e2b for local code completion on my local machine.

    This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments

  • > The reality is that an incredibly small minority of companies in the world do any real training or optimisation.

    At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.

  • > With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.

    Paul, I think, is talking about achieving outsized outcomes in relatively shorter timeframes (as the timing is just right to be investing in learning this tech) for high agency folks who can also afford the ordeal in wanting to maximize for impact & ambition. Of course, there's real risk one may get no where, but even in failure, given you were building the LLM yourself, you might end up with other adjacent, high reward opportunities.

  • I think companies of all sizes will want their own models, or at least customised ones, for their own specific use cases or competition and security issues.

    1. Both training and optimisation will get significantly cheaper and easier quickly.

    2. Politics will probably get even more insane before a potential reprieve on the 20th of Jan 2029.

    3. The big AI firms will become part of the surveillance capitalism network, if they're not already.

    So I think for self-protection a lot of companies will be looking near to medium term AI independence.

    • The argument is sound, but the maths don't math for now, and it's unclear when/if they will.

      For the time being, unless you truly have millions, the outcome from training will be very net negative, while focusing on building on top of existing AI will yield amazing things if you apply the same talent and effort.

      When it does get cheaper, then it will be easier to acquire the skills and experience too, and the struggle you went through by trying to do it now will be somewhat wasted.

      Besides, I am well versed in this field, and it is not rocket science. There are plenty of software engineering domains that are a lot more challenging, like high-end graphics, large-scale data engineering or kernel programming. People will learn to train LLMs when people want them to.

    • Right, just like companies don't use SAAS.

      In reality, enterprises are happy to offload even risky tasks to others as long as they get some contractual guarantees about their data. Would they like more choice in who to buy from? Yes, but not enough to in-house such a specific discipline.

    • The cost of training a model from scratch is going to be cost prohibitive for the vast majority of companies (even if renting the hardware needed for the 1-2 month training time). It's an interesting learning exercise, and some of the things learned can be applied to other parts of the process. There's also the issue of needing a huge amount of data needed to get decent weights.

      Fine-tuning a model or LoRA based on the companies data set is more feasible but you're likely going to need several runs as you test/try out different base models, parameters, etc. This is why there are a lot of fine-tuned models on huggingface based on base or instruction-trained models from the larger AI companies that have released open weight models (Microsoft, Google, IBM, Mistral, DeepSeek, Qwen, etc.).

      Training is limited on memory first (storing training data and weights) and computation second. Realistically you need to own or rent 2-8 H100/B100 devices or Google's TPUs.

      The majority of workflows for a company providing AI capabilities are likely best solved by tailoring a system prompt for the chosen model, evaluating the prompt and model with tools like promptfoo, and then running it on a compute cloud provider (including AWS Bedrock). If the company is big/financially well off enough they could look at buying the hardware needed to run it on their own servers.

      For other uses like agentic software development you'd need to spin up a suitable model on a compute cloud provider (or local hardware if the model is small enough) and then tell your IDE/editor to use that model. You would need some way of benchmarking and evaluating the models to see if they are capable of doing the tasks you need. -- There have been some tests done by people on YouTube that suggests that Qwen 3.8 27B is a decent model, but your needs may vary.

      1 reply →

    • Most companies that build physical goods don't care for one second about their IT department other than how much money they can save per month, starting by outsourcing whole of it, thus they have little use for internal LLMs.

      1 reply →

  • Did you read his comments on this? It's not to actually do LLM research stuff, it's to trigger and unlock ideas.

  • I think his point is, if LLMs are the future (like computing is the future in the 90s), you should be an LLM-expert (equivalent of becoming a software developer).

    I can see the point. It's unlikely that a 2.4T LLM will be integrated into, say, a pesticide drone. You'll still need some kind of LLMs to achieve maneuvers that "normal" programmings can't achieve.

    But what if everything basically turn into that? Essentially, instead of build me a web app to solve X and do Y, build me an LLM to serve X and do Y. (unless the current LLMs are able to do it end-to-end but then they can hardly write coherent software/personal opinion).

  • This is like saying, “teens shouldn’t learn how to make their own game engine because no one is hiring for that.”

    You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.

    • > You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.

      Writing your own game engine makes you realize that the Unity engine is not really that well written....

      1 reply →

    • Except knowing how LLMs work don’t actually provide much understanding for using them. People don’t use LLMs the way we’ve built on most other tons or platforms. It’s more learning Unity hoping to be a better gamer.

      2 replies →

  • That's not true, because everyone, everyone, everyone seems to want to do training. Which results in a 50 person company training, say, a voice model that then fails, because it's just not good enough.

    In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human)

    ... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels.

    e.g. https://valueaddvc.com/blog/inference-chips-vs-training-chip... (I get it, not great first source, but st

    • The big question is whether companies hold enough proprietary data to do useful things that for e.g. Anthropic, etc. can't easily replicate.

      For some very niche cases I think this is probably the case but for the vast majority, the company's data isn't as useful as they think it is or anywhere near the size needed.

    • Everyone says they want to do training, because it's sexy and an easy way to justify raising mad funding rounds. Some manage, most don't.

      I don't know where you are located, but in EU, in China, and yes even in Silicon Valley, the vast majority of companies do not do any real AI engineering. There's nothing wrong with it, it's just not a smart path for most purposes. You can do amazing things without training, and if you try to train, you cannot get anything amazing unless you burn millions.

      Very few people can afford to play the long game and cross that dessert. And, sure, you will not get far without good engineering, but good engineering is definitely not sufficient and is not the primary bottleneck.

      1 reply →

A lot of people here are responding to the message but not to the meaning.

It would be a good idea for young people to deeply know how these programs work. Not so that they can spend their career building them, but so that they can approach the next class of problems we'll all start trying to solve, with intuition all the way down to the weights and underlying mathematics. And also, to develop a healthy intuition of when "Just LLM it" will not be the right choice.

"Build an OS" wasn't a common university project because we were all expected to go out and work on Windows, but because understanding the bare-metal firmware for a computer helps you deeply understand how to intuit building for a whole class of problems.

  • I'm not sure it's possible to have intuition about systems that work in thousands of orthogonal dimensions. In fact I'm pretty sure most of the research is people trying fairly arbitrary things and testing them and then post rationalising implied understanding of what is really happening on top of good outcomes.

    • I disagree. It's absolutely possible to develop an intuition about extremely complex mathematical ideas, including llms or high dimensional systems. Learning to build an llm is a great way to start building that intuition.

      1 reply →

    • I'd have to push back, though not on the part you'd expect. Your description of human researchers is roughly right: a lot of the field is try-things-and-narrativize-after.

      But the load-bearing assumption is that intuition has to be human-shaped intuition. Humans can't intuit thousands of orthogonal directions because we project everything down into a 3D metaphor and hope it holds. That's a fact about our hardware, not about the systems.

      And the reason why is the most interesting part: nothing requires the compression step. A model or an agent can operate over the actual objects, holding thousands of runs and ablations in context and noticing regularities in the native dimensionality, without translating them into a picture of a ball rolling down a hill. No bottleneck at "can you visualize it."

      So the narrower claim: it's not that intuition here is impossible full-stop, it's that human intuition is unreliable. Your post-hoc rationalization point is evidence for that, not against it. The story exists because a person needs something to hold in their head. Drop that requirement and the failure mode goes with it.

  • > A lot of people here are responding to the message but not to the meaning.

    Well it is framed as quite specific advice.

    (I'm done with mining PG tweets for meaning)

  • > understanding the bare-metal firmware for a computer

    IMO this is still relevant, everything surrounding the LLMs needs such a vast infrastructure that I don't know if I would find it more useful to learn the maths behind ML than CS

    • As a 17 year old, I agree with this. Ofc I'm against all the hate directed at PG, I believe that all knowledge has value regardless of its economic utility, but I understand where the hate is coming from. Personally, I find LLMs boring for now, and I'm more focused on CS and electrical engineering.

  • > "Build an OS" wasn't a common university project because we were all expected to go […] helps you deeply understand how to intuit building for a whole class of problems.

    Sure, but it was for a specific degree with a syllabus that taught you the foundational knowledge. It was not expected from the law students to learn how to build one.

  • Depends, the assumption things are predictable always negatively affects both Market Bears and Bulls alike.

    Indeed, if credulous folks look to the world expecting people to bestow success upon them... than the disillusionment with reality will hit their savings harder.

    The Shrek movie market correction correlations are undeniably funny, and a new film is due July 2027. OpenAI may be going public in the next few months while still losing $2.25 for every $1 of customer revenue, and with 6 other firms sharing over $4Tn in debt disclosed to investors in a footnote.

    There is only one direction things can go at the Peak of inflated expectations. Popcorn ready. =3

    https://en.wikipedia.org/wiki/Gartner_hype_cycle

  • > It would be a good idea for young people to deeply know how these programs work.

    It would be a good idea for _everyone in the industry_ to deeply know how LLM training, inference and "agents" work, not least because it removes the ability of shysters to bamboozle with bullshit.

    But, as much as a good idea it is for the young to understand this, it's the elderly who will be really taken advantage of if they do not keep up - just look at Facebook for good examples of why.

While knowledge is always great, I would encourage people not to seek advice from successful people like this (survivorship bias).

Moreover I am not sure it is even good advice? Would you advise a 17 y.o. to learn how transistors work or how to code (i.e. is LLM training the right level in the stack)? LLM training, a discipline where relevant work is already out of reach for 99.999% of budgets really as essential as this post implies?

  • As someone who teaches AI at a university, understanding the basics of how the LLM works is really invaluable to understanding where and how they'll be applied.

    You're right that it's probably a little too in the weeds, but it's also a nice clear and fun objective that teaches you the basics. Like building a TODO list in JavaScript to learn webdev or a Gameboy emulator in C++ to learn how a CPU works.

  • > I would encourage people not to seek advice from successful people like this (survivorship bias).

    Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.

    What's the alternative really, seek advice from unsuccessful people? That seems worse :)

    Personally I do both, read about what worked for people, also read about what didn't work for people, then ignore both and do whatever the fuck I want.

    • It's best not to take advice on direction of careers from anyone. It's better to find and work on things that interest you, and then take advice from people that are amazing in that specific field. In mid 2000s in Australia all the "top people" were telling me not to get into a software engineering career because it was dead. It's certainly challenged right now, but it took off during those 15+ years.

      3 replies →

    • >What's the alternative really, seek advice from unsuccessful people?

      Seek advice from the averagely successful people, since that is statistically what you're most likely to be.

    • > seek advice from unsuccessful people

      Intuitively, I would guess that they have a better grasp of what made them fail than successful people have of what made them succeed.

      8 replies →

    • Kids don’t even know what it is, even less actually feel what it is. At 17 you think that it won’t hit you, that you will be the one to survive until you don’t.

    • > What's the alternative really, seek advice from unsuccessful people? That seems worse :)

      Learning from their mistakes (where it did go wrong) can be as valuable as.

    • > What's the alternative really, seek advice from unsuccessful people?

      Do both. Get advice from successful and unsuccessful people and take the diff

    • I would seek advice from people who have a theory of why or why not they were successful. A lot of those results were happening in very specific contexts and usually should not be regarded as a blueprint, but as inspiration to whatever I do.

    • The unsuccessful person is much more likely to be actually connected to the daily needs and struggles of a normal person. Paul Graham hasn't seen the inside of a grocery store in the past twenty years.

      I know who the 17 years old is closest to.

    • >Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.

      Pretty tall ask, especially if this is targeted at 17 year olds...

    • > Personally I do both

      That’s what you get from listening to “successful people”. You get to learn about all the things they tried that failed, then the things that did work on that 24th try, which was successful.

      The “survivorship bias” people always seem to assume that the “survivor” lucked into his fortune on his first try ever, so he can’t have learned anything, so we don’t have to listen to him. But that’s seldom the case.

      I’ve written about this before:

      https://expatsoftware.com/Articles/survivorship-bias.html

      1 reply →

    • > seek advice from unsuccessful people? That seems worse :)

      Not sure why would you think so.

      Inverse reasoning is very powerful, and unsuccessful people can give you plenty of "don't do this mistake", which the survivors would not even think about.

      2 replies →

  • I would encourage young people to be born rich. It's the best time in 100 years to be advantaged. Why waste your potential by having your labor stolen?

  • It cannot be bad advice because learning something even slightly valuable is always good in a vacuum (as you mentioned)

    Since he has not given any reference point it is equally as good advice as „just learn everything slightly valuable“.

    So the real question is: „What should I not learn in favor of learning this.“

    Or in other words: His comment is not usable because it can mean anything or nothing.

  • > Moreover I am not sure it is even good advice?

    I think it is. He isn't saying to learn how to train a LLM so that you can go on to train LLMs. He's saying to learn it so that you gain a deep understanding of how LLMs work. Ordinary startups can still benefit from things like training or fine tuning highly specialised smaller models, knowing how to select and configure an appropriate model for the task at hand, knowing what software to use and why, understanding what's going on behind the scenes instead of treating everything like a black box, having a higher level of intuition about LLMs generally, etc.

    Most computer science courses do in fact teach things which are lower level than coding, such as how transistors work.

  • IMO these kind of advices never matters. Any individual still needs to make tens or hundreds of little decisions (every day) themselves, and that's what really makes all the difference.

  • > Would you advise a 17 y.o. to learn how transistors work or how to code

    how many of us out here are doing work directly in what we got a degree in? I majored in economics and now I'm a CTO.

    I would absolutely advise a 17 yo to learn how to code, understand how transitors work and how to code an llm. even if he never works on llms, you basically end up with a kid with applied knowlege of statistics, math, physics hardware, logic and a whole lot of practice in critical thinking.

  • As a 17 y.o (way back in the last century) I didn’t need to be advised to learn about transistors. I just had a thirst for the knowledge. I would encourage everyone to learn something about transistors. They are one of mankind’s most useful discoveries.

    • Hey, I'm 19 CS undergrad but i don't really know what to do but I really wanna build something that can shape the world. Can you tell me more about why you would advice someone to learn about transistors? what are the possible career paths?

      2 replies →

  • Right - like advising people to learn nuclear power back in 1992. Perhaps a good idea, but super specific already. My gut is ML/AI is even more complex in 2026… I can’t even remember all the abbreviations and the new ones emerging. And every sub component, such as attention or embeddings, are actually a discipline of its own already.

  • >Would you advise a 17 y.o. to learn how transistors work..

    Hell yeah. Transistors are pretty awesome.

  • I just don't think he realizes how saturated it got over the years. Or maybe he knows at a conscious level, but not subconsciously.

    In 2000 (his era), it would have been really smart to study the source of Linux or Apache. Would have paid dividends over decades. Cuz that knowledge was so rare. The number of people hacking on LLMs now dwarfs the number of people hacking on web servers 30 years ago, by several orders of magnitude.

    And if you turn back the clock even more, I mean just even having access to a computer, let alone owning one, would have put you at a massive advantage.

    I don't know what to call it. The pioneers should be respected obviously, but at the same time you need to understand that for them, the game wasn't nearly as played out as it is now.

    I just don't think you can afford to be dicking around with LLMs like you could afford to dick around with random Linux distros 20 years ago. Too many people willing to do it for free these days.

    You don't wanna end up being the 2030 equivalent of a certain SNES emulator developer, or maintainer of a package manager for jailbroken iPhones, I mean the list goes on and on. Being a hacker doesn't automatically give you a path to being rich, or even making a decent living. It hasn't been that way for a while.

I'm (more than) twice that age, but I've spent time learning this exactly this from videos by Andrej Karparthy and from books by Sebastian Raschka.

I didn't do it because it was useful to me in a practical sense. It's because LLMs are fascinating and I want to know how they work. From that perspective it's been a great experience. I have afirm grasp of the basics. This makes it much easier to understand frontier concepts like compressed latent attention. I can follow the field and understand it.

Not sure I would have got as much out of it at seventeen. I have a lot of background and experience which made it much easier to learn. I wasn't struggling with the linear algebra or with python. I already knew pytorch and neural networks. That helped a lot and I covered these tutorials fast and could skip over large sections. A few evenings and the odd weekend day over a couple of months was enough for me.

For seventeen year olds the tutorials are good enough to make it possible to learn this but it would have taken a lot longer to understand. On the other hand I would have learned a lot more. I think I would have learned a lot of valuable stuff.

However I also think 17 year old me was studying for his A levels and probably this was right choice in terms of maximising future opportunities. I'm not sure I think learning about LLMs instead is sensible. Indeed it might be bad advice. But I can absolutely agree with the sentiment.I think 17 year old me would have wanted to do this too.

  • Which resources from these two would you recommend? Or just blanket-recommend all their videos/books?

    • I kind of want to blanket recommend but that's not very helpful

      I would suggest starting with Andrej Karpathy's YouTube video: https://youtu.be/kCc8FmEb1nY?is=oiDsrBYJg_MUUmoD

      This video is excellent. I'm a huge fan. Also the video is zero commitment and instantly available which makes it a good way to check you are interested.

      The book by Sebastian Raschka is slightly less accessible but very reasonably priced and the experience of working through a book is a lot nicer than skipping back and forth in a video (for me). Sebastian's blog posts on recent architectures are absolutely great too.

    • I've read Raschka's "Build Your Own LLM from Scratch" book and really enjoyed it. I haven't tried the code yet, but the code from his previous Python ML book worked great.

    • actually just use chatgpt. There is a new mode of learning thats now avaiable that doesnt require you to read about things that are already discovered leaving you with a shallow knowledge.

      you can now play the inventor and start with question "i want to build next token prediction software" and go as far as you can with your current knowledge while brainstroming with chatgpt as a rubber duck.

      dont not start with a course on probabality , linear algebra or calculus . do not watch 3 hr videos or 3blue animations .

      there is agood video on this way to learn.

      https://www.youtube.com/watch?v=cbiyPOn-__M&t=380s

I appreciate the sentiment and I'm pretty curious how I could train an LLM, even are really basic one from 5yrs ago without nuts hardware.

That said, I a good starting point for a 17yo is reading about perceptrons[0], then the basics of neural networks[1] (ex. 3-layer perceptron) then writing a program to train a 3-layer perceptron and classifying the MNIST dataset[2] - a dataset of characters.

This can anywhere between a day and a week and you will demystify the basics of neural networks and work your way forward with more advanced contemporary concepts.

Fun fact: any multi-layer perceptron neural net can be reduced to a 3-layer perceptron network.

[0] https://en.wikipedia.org/wiki/Perceptron [1] http://geeksforgeeks.org/deep-learning/neural-networks-a-beg... [2] https://www.kaggle.com/datasets/hojjatk/mnist-dataset/data

I'll use this post as a shameless opportunity to tell more people about a little side project, I made:

http://languagemodelbuilder.com teaches you (in a few hours to days) how to build an LLM from scratch. It's entirely free, without accounts, and without data collection.

  • This is the coolest link I saw on HN the last 2 months. You rock man !

    Thanks a ton for building this.

  • I wish this was available for young folk in Romania: As it happens: macOS: 2.70% of desktop operating-system usage in Romania; OS X: 2.55%; Combined Apple desktop share: approximately 5.25%; Windows: 90.65%; Linux: 4.02% or so AI tells me.

I am kind of amazed how negative the comments are here, especially on HN.

Learning to hack something together in high school using the latest technology (vacuum tubes, radios, microprocessors, web/javascript) has been a common theme in the tech world for generations. With LLMs and online tutorials, this isn't even a difficult suggestion. Do people think learning new tech is somehow wasted effort?

  • If he has said "to learn the math and programming skills needed to understand how to build LLMs" it'd have been much more positively received.

  • > I am kind of amazed how negative the comments are here, especially on HN.

    I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.

    Hazarding a guess, I think the moment Altman became the CEO and later during COVID, the sentiment seemed to have been shifting towards what we see today. But this is all based on hazy memory, rather than looking at the data. I'm sure there is a blog post waiting to be written about analyzing the sentiment of comments to PGs articles on HN, and you'll see a shift somewhere.

    • Imo it's a breakdown of trust of the startup ecosystem as a whole. Repeatedly startups have enshittified and it's become undeniable that the investment apparatus around startups is partly responsible. We have seen a great driver of uncreative destruction, industries undermined, small businesses undermined just to drive masses of money into few pockets - less fairness for the people working in what is now the gig economy and ultimately prices and other costs that end up as high or higher than they were before for consumers. Not to mention the whole AI/OpenAI situation which many perceive as threatening their skillset per se, essentially tearing up the social contract that existed on this site.

      The sycophancy on HN is starting to break down because there is a higher proportion of users sceptical towards the outputs of the VC and wider investment world than ones who believe they're potential beneficiaries of it.

      Tech industry people are becoming less interested in HN as a warm handshake into the startup world because, frequently, they're disgusted by it. And this reflects on the sentiments people post on PG's articles.

      Increasingly if those at YC want the same kind of low-bar praise they got before, they will need to get it from machines.

      1 reply →

    • Because at some point in life everyone gets tired of fairytales. He started mending the anecdotes to his content which always rubs people the wrong way.

      2 replies →

    • > I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.

      Hard disagree. This submission is still being highly upvoted, while another recent post[1] on the harms caused by Graham’s fellows[2], with a fairly tame comment section, has been flagged. That is a constant on HN. It’s not a fluke, it’s as predictable as the sunrise and getting more pronounced.

      I’m sure we’re both biased in our perceptions. Mine is that HN in general (certainly more than any other website) used to worship[3] everything he wrote, together with others like Musk, until things started to really go to shit and many eyes have been opened to the effects of the unfettered greed of rich tech guys out of touch with reality.[4]

      [1]: https://news.ycombinator.com/item?id=49411762

      [2]: A better English word is escaping me.

      [3]: That word I choose hyperbolically but deliberately. It definitely was not “interesting and maybe true in some ways”, it was much more hardcore than that.

      [4]: That is not “knee-jerk” but a slow realisation still ongoing.

      4 replies →

  • I think it is more about people are a bit sick of filthy rich people giving this kind of advice. I would also not read anything he preaches.

  • Completely agreed. The point is the knowledge, the learning and the journey. If a kid has a passion for building or toying with LLMs, then of course, by all means, please start tearing them apart or even build and train your own model. You'll learn a ton, even if you won't necessarily end up using it here and now. The learning experience will compound and of course that will be useful.

    The above is, after all, the whole genesis of the word 'hacker'. We should celebrate that.

    • How exactly does one go about "tinkering" with an LLM? Any architectural change you introduce needs fine tuning. That needs data and compute

      I tried to modify the embedding output of bert to make it generate box embeddings instead of point ones. At the time I had access to university provided A100 gpus but even with all that a training run took half a day. Models these days I don't think I can train it in any reasonable time with that much compute.

      1 reply →

  • Now I'm curious, do people actually tried to hack vacuum tubes or other big servers that's barely 1MB RAM? It seems like another thing that needs big investment to work properly, unlike those other techs where results can be shown even with little materials.

    • > Now I'm curious, do people actually tried to hack vacuum tubes or other big servers that's barely 1MB RAM?

      "Barely?"

      It's insane to lump vacuum tubes together with servers with 1 MB RAM. My first PC, which I used for a decade, had 640KB RAM. And that was an upgrade from 512 KB RAM. My other PC had only 128KB RAM. None of these were considered the equivalent of (by then long dead) vacuum tubes in their day.

      You can get a lot done in 1 MB.

    • Yeah, there are retrocomputing hobbyists who mess around with sometimes-physically-large computers that were important many decades ago. I don't know if anyone is hacking on vacuum tubes of the kind that you could in principle build a computer with - there's a reason they became obsolete for digital computation almost as soon as the transistor was invented. On the other hand, I personally think it would be neat to try to build a CRT in a garage, which is of course a type of vacuum tube. I don't think this would be an easy garage project, but it does seem like might be tractable for someone who understand physical manufacturing and electronics well, has access to glassblowing equipment, etc.

      1 reply →

  • > Do people think learning new tech is somehow wasted effort?

    No. But funnily enough that is a promise by some of the AI cretins and their boosters. Oh yeah best case scenario you learn how to build LLMs for us. We’ll employ you. And then ultimately that just becomes training data for the LLMs to do it themselves.

    But why are people cynical? they ask.

  • > With LLMs and online tutorials, this isn't even a difficult suggestion.

    Don't many of the commercial ones prevent you from using them to build LLMs?

    I would say the reason for the negativity is not because it's a bad idea for a project, or that doing projects in general is a bad idea (it's not!), it's because it's a very specific thing that is not for everyone. The best thing about computing is the low barriers to entry. You can basically work on anything that takes your fancy. So those who are interested in ML will be drawn to learn about LLMs. They don't need anyone to tell them to do it. Telling everyone to do it reminds me of the "just learn to code" stuff of a decade ago. No, please don't, please find something you enjoy.

i was trying to build llms from scratch at 17. failed miserably because i did not know linear algebra. ended up in a different but adjacent field. when chatgtp got big suddenly there were so many people doing llms and i didnt want to compete like that. so much of what was happening was hype, and that really turned me off. am back to building llms from scratch, but like, its a journey teaching myself all the theory on top of my job. i have decent fundamentals, but i need a better grasp of all the advancements in the field in the past 5 years before i would feel comfortable designing anything. baby steps, essentially. am working on better understanding all the layers of an ai while implementing a rag on my local model as an experiment. 17 year olds should learn whatever theyre interested in but need fundamentals in order to do anything advanced.

If I were 61 and wealthy, I wouldn’t write a stupid shit like this.

I’m no paulg, but if you’re reading this - and you’re 17 - just focus on getting into a university and having a good time that you won’t regret later. Play games/sports, make relationships, fall in love, explore.

  • Is going to university really that good of advice nowadays?

    Everybody goes to college nowadays and the average white collar has lots of debt and relatively minor financial benefits over a skilled trade worker.

    edit: woah, so many people insulted by that. In my bubble and friends, me and another friend are the only people that make very good money compared to non-graduates. Plenty of others opened their shops, went into trades, one learned to tattoo fake eyelashes, one became a (successful) farmer and most make significantly more than the average law/chemist/mathematician/physics/architecture/languages graduates. Sure, the lowest salaries are to be found among the non-graduates too, but I don't see any evidence that graduates make that much more, and that graduating is worth it.

    Some answers talking about how "formative college is", but my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.

    • Depends on where you live (lots of countries have no tuition fees or far lower than the US), what funding you have, how good a university you go it, what you want to do (some careers require a degree), whether you will enjoy it, and whether you are there just for financial benefits or more than that.

      Its not good advice for everybody, but it is good advice for a lot of people. What if you want to be a doctor? What if you want to work in R & D? Not everyone enjoys working in a shop or a farm. Also, how old is your friend group? If they are mid twenties you are ignoring the greater scope for advancement in a lot of white collar careers.

      > my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.

      Within the narrow limits relevant to her business. How much does she know about macro-economics or financial economics, or scaling up a business? I also suspect you are comparing her to people who went straight on from bachelors to MBA (which is a bad path - study business after having some experience IMO) and lack experience. How will she compare in 10 years time when those people also have real world experience?

    • If you’re talking purely financial benefit: yes, it is still good advice, as most white collar jobs still require a degree.

      If you’re talking about a place to mature, around others who are at a similar phase of life, also yes.

      It is where most people meet their cofounders, for example, even if they don’t found anything until much later.

      3 replies →

    • It depends on the university and your goals, I suppose.

      Many people are surprised to learn how affordable elite colleges are if you genuinely need financial aid. I had no idea -- was pleasantly surprised when my alma mater took over 80% off of my tuition.

      1 reply →

    • > Is going to university really that good of advice nowadays?

      Just some anecdata, but every single one of my university professors was quite bad, but they think that since they are the professor, that means they are smart and the expert.

  • So get into debt and spend money is the advice?

    Then vote for someone who will make the debts go away?

    • As opposed to what, neglecting the human experience to grind yourself to the bone for those who own capital, and then voting to uphold that capital? Makes no sense.

    • > Then vote for someone who will make the debts go away?

      As opposed to vote for someone who gives even more power and money to the oligarchs? You bet your last dollar that people will choose the former over the latter.

    • yes?

      I mean, have you seen the options for people graduating right now? How people are behaving?

      Or forget the data, look at how the story of the new future technology is being told. The people making it recognize that it has the potential to put swathes of white collar workers out of jobs, and they are openly talking/warning/PR-ing about it.

      People in tech and SV, the places which have a underlying culture of near delusional optimism, are talking about trying to avoid being part of "the permanent underclass".

      Gambling is up, and prediction markets are being treated as financial investments. Wall street bets is a thing, and outright speculative investments are the hope people have to get ahead.

      This is happening in the USA, forget the weaker or smaller economies.

      When people see the future as one massive zero sum game, with no way to win by building, then they are going to change how they plan their future.

Horrible advice. This may have been good advice 10 years ago, but not today. There are no positions for people who "kind of understand how toy LLMs work" because so many engineers do these days. Most of the real LLM optimization work is at the edge of research and highly proprietary and not something you could ever do without infra that costs millions.

But of course, 10 years ago this wasn't obvious.

  • you don't reach cutting edge immediately. you start with the basics

    • I think the answer then is to get a Phd? I mean, one should learn it to satisfy their curiosity and to build things but I don't think it's going to help you career in any significant way.

  • What would be better advice for a 17 year old?

    • Throw the computer and the smartphone out of the window.

      Or more reasonably, the same old thing : use Linux, hack a little, why not learn programming basics. But learn to own your technology, fight against centralization of technology. The same old RMS story.

      IDK where the tech industry is going, if there will be jobs anymore or not, but what I'm sure (and what have been the case for the last 10-15 years anyway) is that for most tech jobs, having good technical knowledge beyond the basics is pretty useless and will probably not be recognized.

      If you can, stay a computer geek if that's your thing, but don't make it your career choice, the Eldorado is behind us.

    • "it doesn't matter what you choose to study because no job is safe from being outsourced overseas, handed over to an indentured servant, or made redundant by a machine. you will spend your whole life surviving while the cannibalistic pedophiles who own everything invent new ways to make you own nothing. be frugal, don't get married, don't have children, do everything you can to stay healthy and independent, and you just might live a reasonably comfortable life."

      4 replies →

  • Yeah he's a moron with a lot of money, that's about it. I'm sure he's said the same thing about various other bags he had bets on throughout the years.

I think a problem a lot of people are grappling with here is that due to LLMs and AI generally, it’s basically impossible to predict what the future will look like or what jobs will still be around.

I’d probably say something like: do something you enjoy and seems like it might be useful, but accept that the pace of change may mean that whatever you study ends up being irrelevant.

Whatever solution there ends up being to this, it’s not going to be one that an individual 17 year old can implement. We’re past the point where individual good and bad choices matter that much to economic outcomes.

  • I would advise any somewhat ambitious 17 year old to avoid tech and get into healthcare if they can stomach human interactions and bodily fluids. Sure, it is not all sunshine and rainbows, but there will still be plenty of work helping people who are ill or elderly. Even in the worst-case economic scenario, medicine will be a more socially rewarding and stable life path.

    • This seems like a much more interesting question to me.

      Telling other people's children what to do is easy and basically doesn't have any downside to being wrong. With your own children things are a bit different.

      So: what are people here with school age children telling their own kids about the future? If their kids ask, what kind of careers would they encourage them to pursue, assuming they have the skills and interest?

      When I was last in the Bay Area, maybe about a decade ago the bookshops were full of titles like "Python for Preschoolers" (I exaggerate, but only slightly). Clearly at the time a lot of people working in tech thought that cultivating an interest in programming was going to be the path to being a successful (by some metric) adult. Is that still the case?

      1 reply →

    • The real reason to recommend this is that they have an excellent chance of avoiding the "nerd-to-incel" pipeline that tech all-but-guarantees for its best nerds. Post GenAI boom, the social cost of working in tech combined with the coming collapse in high paying jobs, means that unironically coders should be learning a bit about coal mines.

      In this regard, Healthcare is a polar opposite. It's pretty hard as a male nurse to not "accidentally" become a home wrecker.

I built from scratch a sparse text embedding model trained on a 13T token corpus. Not the same as an LLM because there’s no transformer in the mix, but still I learned a shit load of things in order to solve all sorts of problems that emerge when you try to access big datasets and daily update tables with billions of rows. But if it wasn’t for a specific use case that I tried to solve I don’t think that whatever knowledge I gained could be utilized in the market. Sparse models are a very small niche and most people I’ve come across with similar knowledge are in academic circles, not business related ones. So even if LLMs are all the rage these days I doubt the demand for people who know how to build them is that high. Someone who knows how to setup an open weight model and expose an API might be more valuable to a company these days.

17 is such a fantastic age to be free and experience the world, you won't get that much of an advantage as these FOMO groomers are selling you into if you start now versus later.

If you are 17, go be yourself, whatever that is, in whatever way you want that to be, but do it so authentically and fully. Be unapologetic about what you love and what motivates you, and pursue that with passion and commitment.

Yeah, no way.

I'd move to the middle of nowhere and work multiple jobs on a farm and in construction. Learn how to grow food, and build things. Meet the farmer's daughter, and marry her. Then, buy my own land, grow my own food, and build my own things.

  • Slight problem with step “buy my own land”. Your working 2 blue collar jobs and paying rent are incompatible with this plan.

  • I already do this. I live in a small village where there isn't even a wired Internet connection (wireless only). I work full-time with LLMs, on own product ideas and client projects (all LLM led).

    I started investing in farms, have 50 pigs and 100+ chickens now. We are planning to grow to 100 pigs and 2000 chickens in a year. We will start growing Shiitake mushrooms in a few months too.

I get this is basically advice for young founders and entrepreneurs, but i would ignore that request and encourage 17 year olds to spend time trying to find a happy medium between work and life.

Being a super rich and an unhappy workaholic, or a super-impressive engineer who wakes up one day at 45 and realizes they regret wasting half their life (I ran into way too many of these) is a much worse fate than "not being rich from your startup" and working a relatively regular job while feeling fulfilled and happy by more than just work.

Especially in the US, which is uniquely bad at this and encourages people to work themselves to death, mental health and work life balance are much more valuable things for 17 year olds to focus on than finding good startup ideas.

In case you think i'm being a bit dramatic, let's look at the state of 17 year old mental health in the heart of Silicon Valley:

"The City of Palo Alto and the Palo Alto Unified School District approved a funded contract to place 24/7 human security guards and monitors at all four local Caltrain grade crossings, including the Churchill Avenue crossing directly adjacent to Palo Alto High School."

(in case it's not obvious, it's because of suicides by high school students)

The 17 year olds do not need advice on better startups, and this situation will never get better if we focus our advice on how to be better at work instead of how to be better at life. This will require redirecting the conversations.

  • Thank you. They need human contact, not more "sit in a room alone and get stressed as fuck for little ROI" tech bullshit. Unless the kid has a genuine, self-motivated interest in learning these things (a great, positive thing that should be nurtured), they should file pg's advice under "ok boomer."

I really don't understand Paul's reasoning here. Does he predict more scarcity on the model-level? That layer seems to be almost a commodity now + training is damn expensive.

If you are really 17, my advice is to identify use-case for AI (ideally relevant for businesses) that work most of the time and find ways to make them work pretty much every time. AI reliability is the scarcity right now.

The core point here is that AI is a massive thing (at the moment) so it's probably a good idea to understand it deeply. Not sure why people are so worked up about it.

Paul G is not writing this for a general audience of your run of the mill “engineer” hoping to be employed by someone. He is writing it for future founders. What knowledge / skills you need to develop today to be well positioned to have a startup worthy insight when you are 24.

Most 17 year olds I know don’t want anything to do with AI and see the entire industry as an existential threat.

Nothing wrong with learning the theory and understanding the papers. Getting to that point you’ll have to get your fundamentals down. Might be an interesting exercise.

But as a future? I guess we’ll see. I suspect the next financial apocalypse will determine if there is one. Another AI Winter that may outlast all others so far.

  • > Most 17 year olds I know don’t want anything to do with AI

    The few exceptional individuals will innovate what the millions will benefit from.

    At 17, the mother of Isaac Newton removed Isaac from school and tried to make him a farmer. We all know that wasn't his destiny.

    • There’s no such thing as destiny.

      You could also walk out your front door and get hit be a meteor tomorrow.

If I were 17 again I'd prepare to go for volunteering overseas after high school for 1-2 years (plenty of free options in the EU where you might only need to cover the plane ticket). See the world, you learn a new language, help others and then think about what you want to do.

As someone who's at a similar age and was interested in learning how to do this, there just aren't enough resources to do so. Most LLM research is in the form of academic papers, and there isn't any 'popular' way to learn these things, and besides that all said research assumes you have a B200 cluster ready to go. If you have weaker hardware (say an 8GB nVidia GPU, which is what I have) you're going to be limited to fine tuning small models or torturing yourself working the GPU for days per iteration trying to run things like https://github.com/karpathy/nanochat, which is hardly an educational experience. Renting cloud GPUs is expensive, and at this age the most I could muster up for experimentation is probably $100 or so, which only gets me 25 or so hours on a B200 which just isn't enough. So why would I bother myself with this if I'm already at a disadvantage because of not having access to the right hardware and when surely there are better ways to spend my time? I concluded the only way to learn and be competitive is by finding work at an AI lab somehow (not happening at 17), or studying ML at the right university.

Putting this in contrast with programming, I learned coding when I was 8, and it was incredibly stimulating to learn because you can quickly iterate and there were thousands of books and YouTube tutorials that dumb everything down and teach you fundamentals. All you needed was a $300 computer, and you can learn nearly anything you want, without being gatekept from this or that because you don't have enough vRAM / an sm_100 GPU.

An LLM isn’t hard to make - the training data is hard to get and prepare.

The big companies stole the data. The average person can’t do that

  • Not arguing how big companies got data, but there is plenty of public domain knowledge and content available to someone who wants to get it.

I'm making an assumption here, but I think paulg is implying that "learning LLMs" today is like the equivalent of "learning computers" in the earlier days. We could even divide civilization into two eras: Before Transformers (BT) and After Transformers (AT).

Oddly, I just remembered I did the nearest possible thing to this when I was 17... back in 1991.

On an Amiga, I took various public domain text documents from cover disks and counted the probability of the next word given the previous word. Then spat out random sequences of words from it and printed them out. It was called "Splurge". Basically a very very simple single layer statistical language model.

Some of the sentences were randomly not bad sentences, which seemed amazing at the time!

That kind of thing (and Core Wars and Tierra etc) did lead me to getting a job at an artificial life startup at the end of the decade. But that was in turn about 10/15 years too early (no GPUs).

There's some lesson from this about timing, but honestly I've gained the most as a person when I did something that was fun, ethical and gained an audience. A tricky combination.

I get it the sentiment behind the post…but it has some “let them eat cake” vibes though.

I'm not sure what 17 year old me would have done with YouTube tutorials for everything under the sun available.

I'm much older and less wise now, but I still afforded myself the opportunity to follow karpathy's tutorials to build a LLM from scratch. Got to play with a few ideas. Seen similar ideas turn up in frontier model work, which is quite gratifying.

There are so many ideas to try.

Currently playing with autoencoders that takes A and B and produce latents A', B', and C'. Reconstruction of A is from A' and C', B is from B' and C'

The idea is if C' can be made to improve both outputs, it must store as much information as it can about what is common to both inputs.

If i were 17, I'd try to distinguish who to take advice from, and would definitely learn that VCs have interest to spread a specific agenda in their message. Also, I would get drunk and have as much fun as could, as the misery of working under the treat of being replaced by AI, would simply kill any desire to live past 25.

The only realistic approach is to train on a limited data set which is probably less usable than the comibnation of a custom RAG + one of the many available LLMs.

Also many people/kids don't have access to proper "productive" systems anymore, since the whole computing and electronics industry shifted to make "consumer"-devices like smartphones or laptops made for netflix, gaming and spotify.

Breaking the barrier to build a custom system, install linux (or developer tools for Windows, MacOS) is already a complex AND costly task. It was just way simpler in the late 90s and 00s to get something working.

This essentially comes down to choosing the search for substance (here in the form of technical depth) over short term gratification and quick wins in life.

I think this kind of mindset should be taught way more in school so that people really appreciate learning a subject deeply.

My kids CS teacher asked me what they should do after AP CS.

Previously they had a class where they'd build apps for other teachers. Like tracking when clubs are, etc. But now that's become easy for teachers to vibe code themselves.

I suggested they shouldn't prereq this class on CS. Heck invite anyone in interested in "building things" and they can get practice at building apps for other people / themselves. Maybe that becomes a gateway TO CS - people who want to learn how things work under the hood.

I suggested post CS class for the CS people should probably be building an LLM or something :)

What about doing abliteration, weight pruning, representation engineering, etc, directly to open LLMs instead?

Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.

Modified open models have a very active community around them, without the need to look much further than Hugging Face.

I'd learn a trade in all seriousness.

(Edit: And learn how honest business works)

  • With the hindsight of experience, the remnants of my 18-year old energy go “woah, that’s cool!” at plenty of engineering feats… and my decades-older second brain goes “well d’oh, I could’ve just learned a trade to work on that!”

    I think the last one was seeing a skilled electronics repairman do surgery on a CT machine controller.

  • Depends if you’re 17 with rich parents or not.

    • This only changes whether you are naive enough to believe “honest” business means anything in today’s age. If anything, I worry being honest is holding back smart people who try to compete in a rigged game.

      3 replies →

When I was not 17 at the times of GPT2, I decided to not bother with learning how to build LLMs because it’s too expensive for an individual. This escalated quickly.

A 17yo can train a small GPT this weekend. nanoGPT is a few hundred lines. Understanding why it works is the part that takes a decade.

You can build the program, but to train it is another beast, billions of docs, images, videos, which a mere mortal doesn't have access to

Second in line, build your own agent, that's more in our ballpark, then customize it to your needs, both virtual and physical

I'm usually a fan of pg, but this post is ignorant of modern AI technologies. Building an LLM from scratch is both a trivial and a useless exercise. There's probably in the range of 5000 github repos doing exactly that. What makes LLMs work is scale, and what makes engineering and training LLMs hard is also scale. And scale is not something you can achieve in your garage.

If the goal is to understand LLMs deeply, one would be better served by either joining one of the big AI companies or doing a PhD. And to be honest, I think this journey should have been started 5 years ago, because right now there's too much competition.

I wonder at the “worlds” Mr Graham envisions, and what is their cardinality. Is this the only 17 yro reimagining, or is there an army of 17yro, of which this LLM curious persona is but one?

Why is it important to train LLMs or even fine-tune them? LLMs have proven their point, costs are crashing, and there are more of them than most companies need.

The real value is to unlock meaningful insights and directions from existing data that is there inside companies.

I live far outside any tech city, so maybe I do not understand. But working with LLMs full-time, building for clients and tons of own experiments, I see no value in building on LLMs.

What would you do if you’re 30 years old now?

As a platform engineer being based mainly out of Australia/Hong Kong, opportunities seem to be getting less unless targeting high frequency trading or banking.

It seems like building a startup with the help of some AI tools might be the best bet.

  • > It seems like building a startup with the help of some AI tools might be the best bet.

    I'd only recommend taking that path if you already have your first paying customers standing by or are extraordinarily good at marketing.

telling a 17 year old to get into tech right now is horrible advice, literally telling them to get at the back of a line with a better part of a million more experienced people in it.

  • Telling them not to get into tech is also terribly reactionary advice. The truth of the matter is that we don’t yet know whether tech roles will be eliminated or if they’re just going to follow previous innovation breakthroughs where “one person producing way more work” makes software even more of a desirable industry to be involved in.

    There really isn’t a very strong correlation between tech industry hiring strength and AI as of yet. Various studies that are out there haven’t even witnessed AI workflows contributing more than modest gains in software engineering efficiency. I.e., being able to write code 20-40% faster isn’t a seismic shift in the industry where everyone is getting laid off tomorrow and we’re all replaced by software.

    Even with the questions surrounding the current job market, it’s still an incredibly good ROI career compared to so many other jobs out there.

    For example, in my local area you can get a job as a registered nurse working nights in the emergency room and only make ~$115k.

    I make almost double that telling an LLM what to do from my house in my pajamas during the day with less time spent in university.

    Even if tech roles lose half their salary to automation pressure it’s still a really good gig.

    • “only make $115k”

      This right here is why nobody is shedding tears for the massive employment crisis in tech.

      You make double that shilling ai slopware while they work nights saving lives.

      I would tell a 17 year old that the world will always need nurses, same can’t be said for guys sitting in their pajamas burning tokens.

      What is happening now in tech has been a long time coming, and it can’t happen fast enough.

Easier said than done - where do you get the B300s from?

Better to start working with harnesses, evals, statistical analysis, etc. - where you don't need the huge hardware for pre-training etc.

If you're 17, and seriously curious about how modern day AI works, you might as well just sit down and look at a couple of courses on linear algebra + calculus, machine learning, deep learning, and more LLM specific deep learning. Those courses will teach you how to go from writing your first perceptron to a MVP language model. But also so much more.

I learned HTML when I was 17 in about 1995 and it's certainly taken me on a pretty fun career path. Less technical than LLMs for sure, but 'figure out where the industry is going and move what you're learning to there' is solid advice.

Writing, supervising and training LLMs are now the purview of... even larger LLMs. Optimising CUDA kernels; hand-writing SIMD assembly to speed up data loading; tinkering with your particular brand of DRAM to see if there's anything to gain from optimising for its memory topology and NUMA --- these are now the job of AI.

There is very little reason for humans to get all too engrossed in this type of work now, today, with the hope of being good enough at it to command a high salary in 3-5 years. AI can already do it incredibly well, and they can do it persistently and doggedly 24 hours a day.

  • If this is the job of AI then why does AMD struggle to get the best performance on their hardware? Not enough LLMs?

Can any1 share any resource to learn that skill. I want something that has been tried by you. I can too search on the internet...

LLMs are the new compilers.

I don't think you can really call yourself a developer unless you at least have an idea how to build a more complex software project like a compiler, and maybe have built a toy one either at uni or for fun.

It's not clear how long this LLM age of AI will last (to be replaced by something better), but nowadays any developer should at least understand the basics of ANNs, and more than just the "hello world" of a cat vs dog CNN. An LLM/Transformer is maybe the equivalent of a compiler in that regard - something that we all use and is complex enough to present a bit of a challenge. You should at least understand the basics of how an LLM is built, and maybe building a toy LLM will/should become the new Comp. Sci. degree toy compiler replacement.

Paul Graham:

>"Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch..."

That's funny, Paul Graham, because if I were 17 again,

I'd learn how to program in LISP.

https://www.paulgraham.com/rootsoflisp.html

https://www.paulgraham.com/iflisp.html

https://www.paulgraham.com/hundred.html

(And/or other LISP derived languages... Clojure, Scheme, Racket, TinyScheme, etc.)

I guess "the grass is always greener..." as that old expression, that old "chestnut", goes... :-)

As an aside, for someone interested and who's an absolute beginner, can someone please recommend good resources on how to build LLMs from scratch? Thank you in advance.

  • 1. Build an LLM from Scratch by Sebastian Raschka (https://sebastianraschka.com/llms-from-scratch/)

    2. LLM from 0 to Hero, and nanoGPT by Andrej Karpathy

    • I second 1. I'm a newbie in neural networks and I think it's an excellent book! One of my barriers in ML is the resources, I find them overcomplicated or too simplistic without a mid term. It's not the case of this book, everything is well-explained. Neural Networks aren't fun for me, but this book makes it very interesting.

  • Stanford CS336 is up on youtube from Spring 2026.

    • I think this is the best structured class out there that teaches how to scale LLMs . Hope the 17 year old knows linear algebra. Building an intuition for the shape of the matrices is important. A lot of understanding the 'building from scratch' means understanding choices like why RoPE instead of the original frequency based positional encoding. Start with Karpathy and then go to CS 336

  • Check the front page of this god forsaken website a few times a day and you'll get about 10 different posts a day about it.

It’s probably wise to learn how to build one to understand what you’re dealing with. However, if I were 17 I would lean how to apply an LLM to a problem instead of strictly building one.

certainly better than wasting time with harness and agent workflows that will become irrelevant at the next evolution, same thing happened with 'prompt engineering'

If you’re 17, you might learn hands on knowledge, like tacit knowledge in areas such as lathes, precision engineering, metrology, and other very niche fields. You can also try climbing and explore arts like music, painting, and drawing. And, of course, spend some time in nature.

LLMs are incredibly boring to me as a technology..Not in terms of what it can do, but how it works.

  • Same for me. Neural networks in general.

    I first studied them in 2011, and I was like what? Just a bunch of partial derivatives?

    I keep looking at AI to check if now it's something else but it keeps being gradient descent.

    Okay, it's great that you can perform miracles using gradient descent but that doesn't make it captivating in any way.

I would (and am) going into MLOps. Not just the general infrastructure/systems administration but how to do inference optimization, caching, quantization, memory pinning, vfio passthrough of gpus etc.

Why are people so negative about this? It feels like a fun project and at 17 the stakes are not really high. Something one could easily do on summer break in a couple of weeks.

  • what hardware can a broke teenager get access to in a couple weeks? is this really better than learning how to program?

i am a small fan of pg, nevertheless i find this to be an exceptionally good take and it is strange to me to see so much piling on to this one in particular here.

learning about llms is not useful so that you can make llms later, you want to learn about it so that you can work on next generation architectures. llms before long i imagine will be left in the dust by ebm / physics oriented models especially that can have an embodied understanding of the world. but a lot of things you learn about them are transferable by doing something like this

I'd just build an agent, it's easier than you think and you'd learn a lot about the "magic" of LLMs.

I read a bit about how LLM works, but as a hobbyist it is pretty frustrating that I won’t be building anything useful without throwing a lot of money at it

As if you couldn't learn what LLMs are at any Age.

The core technoology is pretty basic, developing a rudimentary understanding for why the individual parts work as well as they do is tricky.

It’s really disheartening to see how many people don’t know shit about LLMs, by reading the comments… ironic given what OP is trying to say

  • Why is it disheartening? LLMs are a dead-end technology with respect to AI.

    • Have you been living under a rock? We might not get "super-intelligences" (if such a thing even exists) from LLMs but they've proven incredibly useful in basically anything relating to text, including code.

      4 replies →

I've come to accept that some people, when faced with impressive technology, simply want to use it. They genuinely have no interest in understanding how it works. Lately, its even become fashionable to shame people for trying to understand ("you still read code? gross, you know AI can do that for you"..)

I will never understand this mentality, to let yourself be so dependent on something you don't understand at all is to live like a child. But it is very common. I doubt too many 17 year olds will bother even trying to understand what an LLM is, let alone build one from scratch.

... and it would be totally pointless.

I mean first that is already what plenty of 17yo are actually doing, because that is what they do at school or in parascholar activities. There are already countless of such tutorials where you can do that in an afternoon.

The pointless part though is precisely why Amazon and others are hunting for rare books, all the low hanging fruits have been picked already so just training a bigger model will simply mean burning more energy and money. Sure training a small one for the basic principle is a great pedagogical thing, training another one, medium, then maybe a large one, is also good in term of learning the process and architecture, but one should not expect it to be useful out of that context.

Pure players are precisely doing everything they can to corner the market by making their own scale unreachable by others. Smaller players with access to lesser infrastructure are thus betting on different market, e.g. embedded systems.

17yos should definitely build their (L)LMs from scratch and whatever bigger model they can train for free, or for cheap, but they should not expect that to bring them any riches.

  • Why would 17 year old do something that only brings them money? I do not think Mr Graham here is advocating for the path that makes most money as a result of learning how to train a model. I assume that tinkering and learning about LLMs is what enterprising 17 year olds will do to discover ways they can get a competitive edge or further the SotA with their insights further down the line.

    • No wonder the state of everything when 17 year olds are getting pressured to be “enterprising” and “get a competitive edge”. How about learning to be empathetic, respecting your fellow humans, caring for the place you live in, enhancing the lives of others? We shouldn’t be teaching 17 year olds to be greedy, selfish, self-aggrandising blowhards like Zuckerberg, Musk, and Graham. They are not good role models for the future of humanity.

    • he does mention that it would later on be about building a startup, which is about making money

  • It seems obvious to me that no matter what your age right now, you should want to be as broadly educated and intellectually curious as possible.

    You should want to train a LLM from scratch as an intellectual curiosity itch that needs to be scratched.

    The idea of learning one hot skill that has a pot of gold waiting at the end of it was a brief moment in time that came and went.

    When I was 17, we would have said obviously support vector machines are the future. Neural networks overfit and don't work.

2 years before everyone was doing custom training. What happened to all those today when frontier models itself become more powerful than custom trained ones?

All the negative comments here are really missing the point. At 17, you should be building your foundation. Kids that can make a custom CPU, or retrofit an old car with an electric motor, or screw around with nuclear energy... these are kids that are doing it just to see if they can. It's not about jobs, it's about curiosity and stretching limits and seeing what you can do, who you can be.

Didn't his swiss watch essay say he'd essentially leave the industry because there will only be bloat from now on?

What are the best resources to learn how to build LLMs from scratch for 17 year olds?

I have my opinion on this but I'd like to hear the HN opinion, I will just say one thing:

If you are starting with little knowledge, like a 17 year old would, letting an LLM explain it to you is a terrible idea.

  • Andrej Karpathy has a great Youtube series on how to build LLMs from scratch. Perfect for somebody who just learned a lot of high school math. Would start there and get busy with some handson python coding.

I think this is weird advice.

Learn how to make language models from scratch, yes. But learn how to use them, in the context of other machine learning tools, on very small hardware.

When the bubble bursts (and I still tend towards thinking it could burst rather than be deflated in a manageable way), the focus will be on uses of AI that are not like the hyperscalars' products.

People will still be interested in useful AI being added to small things — assistive technologies, home security, garden monitoring, their phones and smartwatches, robotics.

Instead of reductive, reactive make-an-anthropic-competitor advice like this, what about advising 17 year olds to focus on broad, integrated, helpful AI — or on going back through eighty years of history to look at AI projects that failed and reassess them?

I think pg answered the question as “what I’d do as a project” and not “what I’d do as a career.” So the critical comments are kind of missing the point, IMO.

I don’t see why learning how LLMs work is a bad project for a 17 year old.

Optimizing your entire career and the next decade+ of your life on LLMs? Yeah, probably not ideal. It’s almost always a bad idea to make long term decisions based on current trendy things.

And since everyone is using this topic to give their ideal advice to 17 year olds, my advice as a mid-30s guy: seriously consider becoming highly skilled at a specific thing, and don’t be scared off by the idea that it’ll take 5-10-15 years to get there.

When you’re 17-25, the timescale of a decade seems infinite. But it’s really not, and a decade spent “exploring and keeping your options open” sometimes just ends up with you being pretty decent but not amazing at a lot of random things.

Sometimes I wish I had just become a carpenter, chef, electrician, etc. – a specific skill set that leads to mastery over time, rather than the endless exciting-new-thing hamster wheel of working in tech.

If I was 17 again I would sack off work and study and focus on chasing the opposite sex, without the angst I had at the time. I don't know a single person who regrets having had too much sex when they were young. I would not build an llm, too hard, too expensive to run. Learn how easy life is if your morals allow you to grift money of vcs into your own funds and retire.

Really? because I feel like LLMs are already pretty much a commodity, not to say their won't be advances in LLMs but I don't see the models themselves being all that ripe for disruption the way that the web was and such. I'm guessing chip design and manufacturing processes will be more important than models in the future.

When I was 17 I was building Windows Phone apps, bad decision on my part.

  • I also did, and had written code for/on other mobile devices before, and did write a large number of even very ambitious software for other mobile devices later.

    I do not see any past constructive experience as a waste of time.

  • Incredible counterexample, but oddly relatable. I'd probably have achieved techbro 'post-economic' status earlier if I focused on Android dev instead of the shiny (and new at that time) Xamarin for Windows phones.

    • God I almost invested in Xamarin after Windows Phone got aborted, I did spend a lil time on UWP, but thank god Flutter came out not so long after that. After all these years I learned to stay away from Microsoft tech stack.

This is about as intelligent as say "If I were 17, I'd learn digital electroinics". You would waste your time. Sure, in theory it's useful, in reality it's not that useful.

yeah i loved tech so started learning circuity and soldering. but it was a waste of time i made my living learning how to program web applications.

On first glance, this is good advice and in general I'd give the same for this age group. With age, you'll see more of these scenarios come up and if you have experience and foresight, you can provide direction that may prove fruitful to young people. My son and his friend asked "what should we look into and learn?", about 15 years ago, I told them Python and Java; Python because of versatility and cryptocoins; Java for the long term stability in the job market. I despise Python personally, but I could see its potential and still do; especially for AI. At the end of the day, kids have more time than money and its great experience for them to get their hands dirty and find out what they might be interested in; its a long life.

It seems this is poor advice in that it’s suggesting young people should focus on the current problem as opposed to future problems. Focus on the current problem can result in making some money but it will result in making the incumbents more money, which is not disruptive. Isn’t the goal of radical software startups to maximize disruption?

If the two current bottlenecks, for this LLM madness that could very well be a bubble, are processing capacity and accuracy (a second processing problem) then what comes next? Isn’t that where young people should be looking or are we just giving up on innovation?

I remember when I caught PG on reddit arguing with some guy who'd said something mean about him. He didn't reveal who he was. But you could tell from his history - his first post was from before reddit opened to the public.

Good times.

I wonder if you can dig that out of the historical reddit database. I'd like to see that again. I love how everything is recorded now.

  • > I love how everything is recorded now.

    I mean, except when it's censored. That part is a shame. HEY! Is there a tool that tracks the censored comments? I bet there is. And if there's not... that will be a delicious new project for my agents.

It’s crazy how much survivorship bias gets repackaged as generic advice.

Wait no it’s not, that was always happening.

What’s crazy is that people still believe in it.

I don't think individuals have the resources to build an interesting llm. The l stands for large. You need a dataset too. Llms are only interesting because theyre large

And it's basically a weekend project to put transformers together in a ML library and train it.

The follow up comment,train it to play a game also doesn't make sense? Llms Sony really play games and there are better ml approaches to do that?

Terrible advice. If I were 17, genetics and bio tech at the next frontier, with opportunities to be more than another corporate drone. AI is a lot of bureaucracy and nepotistic who ya know already.

Somewhere in rural America is a 17 year old that doesn't even have working plumbing in their house still.

I'm sure they'll get right on powering up their computer from the hamster wheel, Paul.

My heart goes to all the kids out there that didn't get the fair shake let alone fair access to tech that gets these condescending "learn to code/learn to LLM" bootstrappy talks from rich pricks that don't know what life really can be like for a lot of American kids out there.

If I were 17, I’d learn how to invest and build financial literacy, and plot potential growth of my networth throughout my life, before even thinking about a career. Then smoke a bowl.

I get the sense things have changed a bit since I graduated and there are lot more jobs in AI outside of academia these days, but it's still a very different field from other SWE pursuits, and it's not really accessible to hacker-minded people.

Learning AI isn't like learning HTML in the 90s then expecting to get a job at a tech company building websites. You can't just "learn how to build LLMs" and expect a frontier lab to hire you so I'd argue this is rather bad advise.

Additionally, unlike web development in the 90s you cant really do anything interesting yourself... All of the interesting/useful stuff will require huge amounts of compute and data so there isn't even much point in learning to start your own thing either.

As someone whose built many of NNs from scratch (hand written code, long before the days of LLMs), it's more or less useless knowledge if I wanted to work in a frontier lab or do anything interesting in the field.

I also think anyone thinking about going into a field which is basically a crossover of CompSci and Maths is absolutely insane right now. Even if you think there is a place for CompSci and Maths post LLMs, there's almost no chance anything you learn today will be relevant to the skills required in say 5-10 years.

I would not waste my time with yesterday's fad. The next unicorn generator will be something else.

YC Combinator guy says that with a time machine he would learn to build the currently trillions-valued or whatever technology. Okay.

  • No, that isn't what he said.

    He said, if he is 17 _now_, with no family, no commitments, no pressure to start a career, he'd invest his time to learn to build LLMs from scratch instead of trying to start a company.

    • Okay sure it’s not a “time machine” in that sense. He wouldn’t go back forty years or whatever. And he’s still saying the same thing that I was alluding to.

Venture capitalist suggests everyone to become his future employee, just as he has been (successfully) doing for his whole life.

Can't this guy enjoy being rich in silence? His takes get worse with every passing year.

  • He did get rich by being pretty much the opposite of silent, so I'm guessing you can't just turn off that part, kind of comes with the package ;)

This is always such a nonsense question-answer thing, asking a person who already succeeded what they would do if they were young.

Even worse when they ask themselves.

If I were 17, but have the money I have now he means.

  • The times are different. When I was 17, we had to buy records; a 17yo today can listen to the whole of the available produced music, plus interviews and all other uncommon and related material, for free from the comfort of "here and now".

    Possibilities exist now that did not exist before. Those who do not exploit this are fools.

What the fuck does this guy know about? I'm sure if we went back through similar statements he's said over the years he's said the same thing about various technologies that are no longer relevant. The guy is a talentless hack who larps as a blogger and his only "redeeming" quality is having lots of money.

Owner of Golf Club Company says I should dedicate my life to golf lmfao.

I genuinely don't think telling young people to do anything tech-related is good career advice. We don't even know if entry level roles will ever come back. The situation couldn't be worse for these roles. PG thinks that some random teenager will build a startup and get rich from it, or some shit. Like get real, man. Completely out of touch, tech bro who hasn't worked a real tech job for the last like 20 years... Remind me how many startups succeed again, Paul? What about the market dynamics for LLMs and how one might "secure compute"?

  • These guys (VCs) entire livelihoods rely on tens of thousands of young people throwing away their early years of lives trying to get super rich. Only .01% even see moderate success but PG and co don't care and continue to pump the dream and urge folks to waste their lives to try and make them rich.

I think collectively we should all stop listening to Mr Graham..

He capitalizes on greed and hype but with a soft, sober and thoughtful voice so as to lull you with rationalism and now 20 years of his “disruption” has mostly ruined modern society and a whole generation of techies have been led astray into trying to “change the world” is the world of today (minus the magic technology really any better than 20 years ago?)

- Good for him and his Tech Bros, bad for the rest of society

hackers and painters and kids and ROI and startups and capitalism and the destruction of nature and old guys with money talking like they know better in fascist social networks

  • > fascist social networks

    Sometimes if you want to be heard you must go to the public square, whatever the flags there.

    The rest of the post is unintelligible: add some verbs at least.

The amount of people who missed the point here is absurd. He's advocating for learning about how LLMs work. For the sake of learning. Because no one's going to invent the next thing without at least some understanding of the current thing.

I do not think it is a proper thing to do for 17 y.o., unless they are exceptionally mathematically gifted, as proper understanding of how LLMs are trained requires a good grasp of calculus, understanding modern OS and SDE tools for proper implementation of pipeline etc.

I'd rather simply write another mnist implementation and check if I really like all that AI stuff at first place. Even then, before going into mature-on-the-way-to-dying tech (LLMs) I'd rather focus on fundamentals - good ols linear models, regressions, stat etc.

  • > I do not think it is a proper thing to do for 17 y.o

    If I'd get a buck every time someone said something like this to me when I was in the 13-18 range, I wouldn't have a ton of money, but it's so very annoying when people tell you this.

    Regardless if they're "gifted" or not, regardless if you believe in myths like that or not, let children explore what they want to explore, even if you don't understand what it is or why they want to explore that, just let people explore, regardless of age.

    It was such a terrible experience being a young kid growing up, with so many adults spending hours trying to convince me to stop sitting in front of the computer so much doing whatever; "why are you even trying to learn that stuff, you have to go to school to understand anything of this" and so much other similar trash.

    Sorry, not your fault and I'm borderline trauma-dumping now, but really sad to see this sort of gatekeeping on HN of all places, age is irrelevant to learning ANYTHING, in my humble opinion at least.

    Kids, find anything interesting? Jump into it, ignore what adults tell you, and do whatever you feel like, you'll find your place eventually.

    • Agree with this - started programming through learning scripting in ROBLOX when I was like 12 (this was back like 17 or 18 years ago) and it developed into a life-long passion for software engineering. I am thankful I had people around me (my parents), who were aware enough to realize I wasn't just playing video games and gave me the time I needed on the computer to learn and experiment with programming...

      This also meant that by the time I was actually offered to take a programming class in school (junior year of HS), I had already been able to self-teach myself well beyond what that class was covering, thanks to just working on random projects that scratched an itch I had at the time, looking up anything I didn't know or understand, and internalizing those concepts over time.

      In short though, I definitely agree, young kids and teens (and also, frankly, adults too!) should be encouraged to explore things that they have a passion for, without being told 'you need to go to school for this' or 'you cant understand this at your age'

    • I just voiced my opinion. I just think buiding an LLM from the scratch for 17 y.o. is pointless exercise, advising a teenager to do so is borderline irresponsible, and frankly PG is simply virtue signalling here, as LLMs are still trendy, esp. in his circles.

      There still will be varyy small number of outliers among youngsters who'd be able to extract tremensous value from such an excercise, but for most that'd be _IMO_ waste of of time, with illusion of understanding w/o actually having any.

      5 replies →

  • Why would you tell people that the correct order is to build foundational knowledge before exploring a subject? For some (many?) people, a 'proper' understanding develops _after_ the exploration.

    • > Why would you tell people that the correct order

      Because I can?. JK. Because that was my experience, of someone who is 2.5 older than 17?

      > For some (many?) people, a 'proper' understanding develops _after_ the exploration.

      I am afraid you have a too confrontational attitude here, but I'll answer anyway: because I do not believe you can simply "explore" such complex topics like building an LLMs. You'd simply be unable to build LLM drom scratch, unless you'd call cargo-cult chaining magic numpy incantations you've taken from Karpathy's tutorials "exploring".

      If I were in "exploratory" state of mins, I'd rather go from entirely different side - I'd try playing with LoRA-ing existing small LLMs, such as venerable 2 y.o. Mistral Nemo, to get "feeling" for what training is and how hyperameters influence the process.

      4 replies →

  • > I do not think it is a proper thing to do for 17 y.o., unless they are exceptionally mathematically gifted

    I attempted many projects at a young age that I was absolutely not equipped for. The result of the attempts more often than not left me equipped, every time it left me better off. This is terrible advice.

    • That'd would be a terrible advice if there weren't a plenty of other things "you are not equipped for", but far less daunting both theoretically and practically. Such as, say, convolutional neural networks, or some older ML tech. Or even something totally unrelated to ML.

      Transformers are difficult to understand even to people with strong ML background, let alone a teenager.

      1 reply →

  • >understanding modern OS and SDE tools for proper implementation of pipeline etc.

    Can you provide an example?

    • How would you filter out garbage from your training data, for example? If you are trying to use someone elses corpus, would it be "from the scratch" then?

Another great quote by PG. I've been really enjoying his essays recently - truly a great and curious mind.

If I were 17, I wouldn't be using a social media plattform run by racist neo-fascists...

He bases this decision on all of the experience he has amassed, as a 61 year old man in the tech industry. An actual 17 year old, with 17 years of experience, would not think like this, nor should they.

  • Naturally 17 year olds don't think long-term like this which is why PG's advice is so useful. It gives them a pathway to follow that they likely wouldn't have reasoned otherwise.

    I told my much younger brother when he was 12 what programming was and it'd be a great career. He looked into it and within months was writing CLI games. Eventually releasing his own unity 3d game on steam as a teen.

    Eventually he got into CS and did really well because none of it was scary and new. He parlayed that into role at Meta out of university.

    My point being, 17 year olds have time to learn new skills and guidance can go a long way.

  • Yes. And they almost certainly have a better understanding of their own situation that him. This is not a dig at Paul Graham, the closer anyone is in age, the better they understand what they have to deal with. I'm roughly in the middle between Paul G and the 17 year old, and even though I'm really quite fascinated with zoomer culture and probably come more in touch with it than most (due to relatives in the age range etc.) I realize I have very little idea what it's like to grow up in the world they grow up in.

  • Most. But, that isn't the point.

    Either way, this isn't really advice for 17 year olds. Pg is thinking out loud about the pathways for founders.

  • What about a 19 year old?

    >Whoa. I’m 19 and I trained a 100M language model from scratch. Did a v2 now with a new SFT experiment to see if I can get better results on same size.

This is the ultimate builder’s mindset. Tinkering with the hardest technical problems—even for frivolous things like games—always yields the highest return on curiosity. Time to go back to the fundamentals.