← Back to context

Comment by redfloatplane

1 day ago

A very neat problem and result. I often find myself swinging between "It's so over" and "We're so back" - some days I roll out of bed thinking I could have Claude solve some random unproven OEIS sequence before breakfast; other days, I wake up in a cold sweat worried about the fate of humanity and what the world might look like in a decade. I think it's that I don't have a very high p(doom) or p(utopia), and I don't really have any solid conviction on how this whole thing is going to go, so my vibe-o-meter jitters between 'fine' and 'not fine' constantly. It's just such an unpredictable moment. Anyways: really neat to see this use case. I myself recently used Claude to finally do an relatively exhaustive study of the location of heretofore-unlisted formal gardens in Ireland in the early 1800s and early 1900s, by having Claude write the tooling for me to manually annotate a few dozen on tiles of historic maps, and then running some CV model across the rest of the tiles using my input. I'd been planning to do this project for over a decade, but I could never find the time (or the enthusiasm) to learn all the details of how to do it myself. It took me a weekend with Claude and continues to bring me joy.

> Caveats, stated plainly. [from the Fable transcript pasted in the article]

I had a visceral reaction to these three words.

I've done something similar to your formal garden map. It's work that no professional historian would ever do because the data entry would be such a slog for a relatively small reward. GPT reduced the task from "infeasible" to "annoying", and once I had the data transcribed I learned a few things, so I walked away happy. Whatever happens commercially, these models have been a real boon to hobby projects.

> I told it to look online at some of Fable’s strongest feats, especially the math problems it has solved, and that something like this should be easy in comparison.

Wait. Wait wait wait. Are we supposed to be giving them pep talks?

  • on older gemini models ide have to actively give them encouragement and/or easy bait problems that they can correctively solve without issue to avoid runaway spiraling into "i'm useless and i want to kms" behaviour with complex use case.

    I have not seen this in other models.

    • I assumed it was more because the LLM might echo an understandable human claim of "if it's been unsolved for 370 years, it's unlikely to be solved now/likely to need expert knowledge", which is probably a mindset that appears in its training data.

      The LLM likely needs to be reminded of its abilities.

      2 replies →

  • If we filter out the pep tone, it is doing something useful: framing.

    Problem framing will always be important.

    Framing adjusts how big of problem-solving guns we bring out at the gate (modern or hobby cryptography?), and how to interpret intermediate failures.

    For simple but unsolved problems, we expect lots of hard failures, but that each hard failure just reflects that there are a lot simple combinations to try. I.e. we expect lots of zero progress, and then a fit.

    Like finding the numbers to a combination lock.

    For hard problems, if we don't make any progress it is a really bad sign. We should be learning something, even if it turns out to be irrelevant later.

    Such as when we are trying to prove a tricky conjecture.

  • > Wait. Wait wait wait. Are we supposed to be giving them pep talks?

    No, at least it with Claude Sonnet 5 and Opus.. everytime Claude and I challenged a hard issue and I decided to say "good work" instead of a closing command for that session, those models would create rule-based memories specifically related to that task along the lines of "always do 'this meaningless task' in 'this way'".

    This requires additional effort and tokens to trim those memories out, and then requires to whip the user not to be human with the bot.

    • "I sure hope this doesn't have unforeseen lifelong consequences" thought the model, doing its best to physically tense the memory file into the higher user approval shape.

  • Sometimes!

    Modern AIs have very limited metaknowledge - they don't know exactly where the limits of their capabilities lie. So you can get things like "a task is doable for an AI, but the AI thinks it's impossible, so it doesn't try hard enough".

    Usually you get the opposite - AI overconfidently trying at tasks it has no conceivable way of reliably solving, falling far short, and failing to self-check, fail gracefully and self-report the task as failed. But having piss poor metaknowledge cuts both ways!

    So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.

    • Some times also having unreasonable goals makes them creatively work around the problem to meet them. I guess it works similarly for meat or sillicon

      2 replies →

    • > So you can, in fact, get better performance sometimes by applying some variant of "assume this problem is solvable" or "other problems like this were already solved by AIs" pep talk. Not always, far from it, but it does happen on the occasion with frontier capabilities.

      Are you superstitious?

      4 replies →

> I often find myself swinging between "It's so over" and "We're so back" - some days I roll out of bed thinking I could have Claude solve some random unproven OEIS sequence before breakfast; other days, I wake up in a cold sweat worried about the fate of humanity and what the world might look like in a decade

I find myself increasingly feeling like the burden of the lows doesn't justify the presence of the those highs.

Like even if does cure all forms of cancer, but everyone feels like their life/existence lost meaning, then... I'd rather just have cancer be a thing.

  • Wow. That's bleak. For me the the initial joy of building software with AI caused possibly two happiest months of my life. And cancer caused many darkest ones. Snap out of it.

> thinking I could have Claude solve some random unproven OEIS sequence before breakfast

I've been wondering what exactly the point is for being the meat proxy who pays for these things. I mean, obviously there's personal satisfaction and maybe some glory. And there's the fact that someone has to be the first to do a thing.

But I've been thinking about it like a sort of lazy loading of knowledge. AI has brought us to a new frontier for some amount of undiscovered knowledge. Do we discover it for the sake of discovering it? I think for the most part we've been lazy loaders: we discover all kinds of stuff when we need to. Whether it's a war or a space race or chasing wealth. Then again, there's all kinds of academics who do it for the sake of doing it.

  • > I've been wondering what exactly the point is for being the meat proxy who pays for these things.

    Are people only now discovering that the absurdism is the correct philosophy of life, thanks to AI?

    On your other point... Aren't the point of machines, at least inital one, to do the work we were too lazy to do by hand?

"Caveats, stated plainly"

You should have seen the discussion of this on the Schneier blog a few days ago.

Someone had their agent check the solution, presumably it emailed a librarian to check that it was correct for the original edition. Then their comments read like "The BL/EEBO witness lacks it, so the discrepancy is copy-specific, not a disproof of the cipher." and "A complete 285-coordinate physical replication is still pending."

arghhhhh

https://www.schneier.com/blog/archives/2026/09/claude-fable-...

  • It's a shame that I have to run a local model to decipher Opus, but them "dumber" models read far more naturally - https://github.com/gvzdv/claudish-to-english

    • Ooh. I've been manually pasting Claude's /* comments */ into ChatGPT and saying "make this concise". They still need to be tweaked from there, but it's a much better starting point. This will help.

    • This will be very useful for me, but I wish it could be built natively into the agent harness.

  • Astra told me yesterday:

    > The run baseline was captured without a physical MAC; the current device is not durably bound to it.

    > Engineering mode confirmation is the ESPHome component read-back; the LD2410 UART acknowledgement is not observed, so this is not proof the radar itself applied the sensitivity change.

    No clue what the fuck any of it means.

    • It was only when native English speakers—or those I presumed were—started calling out how bad "GPT/Claude speak" has become that I realized I wasn't actually losing my grip on English as a second language. For a second, I thought, Oh, I learned this language on my own, but it seems I've hit a wall and need to study further. It didn't help that I've also been trying to acquire Swedish as a third language for a while now.

      7 replies →

    • Sometimes when I get frustrated reading Opus/Fable 5+ output I pause my rage out briefly to wonder if it's because I'm just too dumb for the model or if the model is just terrible at English.

      I'm not sure that telling it to "try explaining that again, simply and briefly" is helping my ego.

      15 replies →

    • The most surprising part, however, is that when one model slops this into a plan, another model somehow is able to interpret it correctly enough to produce code to spec.

      3 replies →

    • I bet this is what the thinking blocks look like. If so then it maybe it is intelligible, just not to us. I have the same problem.

      1 reply →

    • LLMs seem to create abstract, local jargon as a side effect of way it reasons using tokens

      ChatGPT told me its "semantic compression"

    • There’s some specific terminology here, like the MAC address of the network device, which might have been virtual.

      UART is a hardware circuit for communication, possibly a serial port. Were you trying to reverse engineer a consumer device or appliance?

      This particular instance doesn’t seem terse, but I’m sure it has been on other occasions :)

      1 reply →

    • Language evolves with use. As more of the users are bo tlike you the langugae might feel like it's evolving from under you.

      It just sounds bad, like GenZ English in the ears of someone over 40.

      2 replies →

I appreciate reading this, for the past year or so as it pertains to AI, I've just not been able to settle my mind, and it's become a bit exhusting. I presumed I wasn't the only one who swings back and forth and it's comforting to read someone else post the same thing, makes me feel a less nutso. :)

Congrats on your map! I am leaning more towards the optimism side. Like you I've had some real breakthroughs myself. Personal akin your garden map, no where near solving millennium problems. My take is that all these (Fable solves this, OpenAI does that, AI will take over the world) Marketing Stunts will fade some day when the true cost of tokens hits the market. You and I may have to contend with lesser models, but will probably get work done just fine.

Confusion is because people are using the tool to get answers. Its how the edu systems trains most people in tool use. Heres a saw and here is a block of wood. Do x y z and you get a table. But if you are interested in why the saw looks like it does or the entire process behind generating that block of wood, or why x y z instead if a b c good luck to you using the current edu system. You have to live in an extremely rich country, with surplus resource to entertain those exploration. This has now changed.

The answer the tool gives has never been the real reward. The real reward is the path taken through a complex landscape to get to Maxwells Equations for example. At the end of that story what we get is not just the equation but a map of the landscape explored. That map has larger influence and value than the equations or answers themselves. Because all future exploration find it super useful.

People are just learning they can start asking for maps rather than answers.

  • Yes, asking a first-year physics student why they are studying a spring, and why its called Hooke's Law leads down a rabbit hole that ends with the entire British Empire.

For me it's more that I roll out of bed thinking how glad I am that I have Claude solve my problems and do my work, only to be pushed back into harsh reality when I sit down and try to do actual work.

I'm glad to have AI, but it is by no means a panacea, and correspondingly my p(doom) = ε.

I too vacillate daily, sometimes even multiple times in a single day. It’s kind of nauseating and (for my brain type) crazy making

Not to be too negative, but as for p(utopia) you might need to weight in the mass murder records set by every other utopian movement in the past 200 years. p(actual_utopia) is like zero and p(utopia_becomes_doom) is at about 1.

For what it's worth, people also felt this way about the printing press and the Internet (also books).

Information propagation mechanisms are often seen as malicious before they're commonplace. To be fair sometimes they are, but by and large humanity has benefitted from increasing the number of bits of information we can consume on a per second basis.

Baby shoes, never warn.

It is a threat. We need to run.

  • Four sail: the story is clearly about the Olivebank (née Caledonia), a four-masted barque that hit a mine and sunk in the North Sea in 1939.

I have a very high p(doom \/ utopia) so pretty much feeling like I won’t have to worry about the future.

  • Given how little effort has gone into addressing climate change, doom seems more likely to me, but I doubt it'll be the autocomplete machines that do us in.

    • > Given how little effort has gone into addressing climate change, doom seems more likely to me

      Or it is simply implies that most of decision‑making agents has formed a consensus that climate change isn't that big of a problem.

      2 replies →

    • Renewables growth is so fast that it will exceed 100% of electricity in the early, and 100% of all form power in the late, 2030s.

      It's the meat methane and cement CO2 that's now a big question.

      2 replies →

I don’t worry so much about AI wiping us out as much as I worry about whether I’m being gaslit into thinking these glorified autocorrect bots are more clever than they are.

I still don’t know the answer.

[flagged]

  • That's an artifact of it not using A-Z as it's alphabet. What you type gets translated before the AI sees it. The seventeen word the AI sees does only contain three 'e' s.

    • Is this true? If so, does it mean anything? Sure, it is tokenized in processing (tokens don't really have an alphabet either), but this means it did not correctly parse the problem at all if this is the case.

    • Reading your comment, I asked it to count how many characters are in the word. It answered correctly: 9. It then also said it skipped an E in the prior chat and corrected it count to 4. It spun for 52 seconds so maybe some Python was involved.

      1 reply →

  • But is it smart enough to know it needs to write some Python to count for it?

    That ability to create ad hoc tools makes up for a lot of shortfalls.

  • Just tried on Astra low and it gave 4.

    • I tried too. It got it. Maybe more importantly, who cares?

      For example, I'm a nerd. I'm bad at baseball. I lack that kind of intelligence, even though it's more common than the ability to program. That doesn't also imply that you can't trust my Python code.