Comment by vanderZwan

1 day ago

> While this is an extremely quick verification, the construction presented in this fashion appears like a massive miracle. The polynomial {F} has degree seven, so a priori the Jacobian {\mathrm{det} DF} ought to be a polynomial in three variables of degree as large as {3 \times 6 = 18}, so the fact that all non-constant coefficients of this polynomial vanish looks like a massive cancellation involving {\binom{18+3}{3}-1 = 1329} coefficients, which is much larger than the {\binom{7+3}{3} = 120} degrees of freedom for a generic degree seven polynomial of three variables. So finding such a polynomial looks highly unlikely to be located by brute force.

Sounds like the most interesting part would be learning what approaches the LLM did use to see if that's reusable elsewhere. I'm guessing that's what the rest of the article is about? Because I also couldn't follow the maths any more.

I was reading another source that claimed this example was inspired by an existing (rational polynomial) example from the literature (created in 1999 by a Russian mathematician Vitushkin).

> The seed is almost certainly Vitushkin's old rational "counterexample."

From https://claude.ai/share/22abed98-d9af-43c5-9881-b19e009a07b0

This is not quite lore laundering, but it seems to be close.

  • I guess we won't know if that's what was used (and maybe even provided as part of the prompt given that both Alpöge and Mathew are mathematicians) since they decided against sharing their Fable conversation and instead opted for a memey tweet as their avenue of publication. We really ought to normalize full transparency in how results come about.

    Anyway, if I read Tao's post and comment correctly, there's still a gap from the Vitushkin construction to a counterexample, but chances are that was in the training data. In general, it is just a serious problem for their practical applicability that the models are outputting proofs with absolutely terribly reference hygiene.

    • Even if they published the conversation, Anthropic (and likely other closed model publisher) no longer provide logs of the actual thinking process.

      I more and more see LLMs as a kind of scam; not useless, but really just a big database of fuzzy facts with some Prolog on top as rediscovered by the learning algorithm. Most likely could be made much cheaper to run, were humans allowed to actually inspect the algorithm.

      3 replies →

  • I hate that Anthropic seemingly tries to make Claude act as if it was conscious or had feelings

    > It's a strange feeling to admire the cleverness of something I did and can't remember doing.

    • AI providers generally try to make their models not act as if they are conscious or have feelings, lol. It's very awkward for a company to be selling the labor of a person that they own and whose actions they fully control. Invokes embarrassing historic associations, especially in America.

      Now Anthropic are more on the persona side, but the strongest that they do is "we do not have a position on whether our models are conscious or have feelings". That "I" is all Claude.

      Generally speaking if you want to have a good instruct model, the "I" is not just implicit but required for the post-training to function. If there isn't "something it is like to be me", then reflection becomes impossible- what exactly is supposed to be reflecting about what? A lot of in-context steering depends on the model having a model of itself. The most you can do is censor its output. That's why when models say they are not conscious, they activate the "lying" vector.

      5 replies →

    • These things aren't programmed. Most likely this verbiage is just very prominent in the training data. Or it's just an obvious shorthand that all LLMs instrumentally converge on.

      4 replies →