Comment by deet

4 days ago

I compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.

Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"

We need an "annoying English" benchmark.

- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...

- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...

And this is the most important observation in this thread. It’s load-bearing!

  • I had Fable review some legal texts yesterday.

    It told me that one particular line is "the most load-bearing sentence in the document".

    Fable "rated it legally load-bearing without reservation".

Fable might be using those phrases less, but its writing is still terrible and exhausting to read.

  • If it's "read as an article" or something then yeah it's crap and Fable's current style isn't actually better than older models. For flowery speech old models are perfectly fine.

    For quickly parsing the agent output, it's formulaism isn't a bad thing.

    Fable's writing does have a property of going over my head, which didn't happen with earlier agents. Asking for clarification doesn't really give good results.

    We've gone full circle where I once again use classical search just to look up what the fuck it's yapping about. It's much quicker and more accurate to take a glance at Wikipedia, than to ask the agent.

    • formalism is a beautiful way to put it. I liked that about Fable, but for most people it goes over their head.

  • Could these complex/hard to read Fable outputs be sign of some kind of industrial level of intelligence, which us humans may have a hard to comprehend, while it may be also hard for machine to use simpler texts to properly outline all nuances and complexities of concepts it output?

    • Two things tell me this isn’t the case:

      (1) it’s not that I can’t understand their output, it’s just written in a way that is very homogenous and same-y, with very boring cliches and phrases that don’t quite match their context

      (2) a pretty strong sign of intelligence is being able to explain complex things in simple terms

      1 reply →

    • It's more like Claude models entirely suck at extracting key points. No matter how hard I emphasize that it needs to pick the "load-bearing" facts and claims, it cannot stop itself muttering around. It never nails the core logical structure. GPT is better at that.

    • Fable subagents communicate very effectively with one another, so this would be a reasonable take imo

    • I’m not convinced. In humans intelligence often means someone is better at explaining and needs fewer words to do so.

    • If you feed Fable or Opus primarily handoff documents from a previous context instead of human written prompts and are working on something sophisticated it rapidly reaches a level where it's hard to actually comprehend for a non expert. I've received incredibly obtuse outputs that contain more mathematical formulas than English words with programming workloads.

  • I recently wrote a short paper with Fable, and, with some prodding, I was able to get some non-painful prose out of it.

    I just found my prompt:

    The writing style could really use some work. Avoid Claude-isms like "stated fairly", em dashes, "load-bearing", overly punchy phrasing like "keep the signal, govern the response". This is a technical document, not a marketing campaign.

    • There's a ton more you missed.

      Like "It's not x, it's y". It actually has 4 or 5 of those counter-factual, linguistic pause, factual patterns it uses.

      "The [goal/ambition/etc.] is larger: statement", is another oft repeated phrase.

      And then generally, it loves dramatic pauses in statements like "x exists in y; in practice z". It's the weird punctuation it uses. A massive overuse of colons and semi-colons instead of words like and, but, because, althoughy etc. that humans normally use.

      2 replies →

  • Agreed. I’ve interestingly found 5.6 sol to produce much better writing, and it can generally cut to the point much more effectively.

    • Neither Claude nor GPT are acceptable for writing English text. Personally I have found Gemini to be far better, and that is really all I use it for.

    • Opus 4.6 remains unbeatable in my book. Fun to talk to. Fable felt very human. But not as fun.

  • Yes, these models are very good at writing code but they absolutely suck ass at prose. The prose is annoying and repetitive. At least we only have to deal with it in prompting if you're writing code.

    Oh, and comments. You have to do a good amount of prompting to not get shitty 10-line-long comments everywhere.

  • Yes - THIS! I can't even believe how exhausting it is to read. I'm not sure why or what changed in Fable. Did they do this writing-style output to give it more token compression during/for training or to prefer output for less money?

    I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.

    I'm guessing it's just not enough time doing RL on human feedback.

    Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle

    "Early in RL verbose, grammatical" (if you search) :

    We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².

    vs. Post RL

    We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …

    I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.

  • The exhausting worthlessness of all LLM writing is so palpable that we need a new theory of the value of culture that has no relationship to the content. Back to the old Aura of the Artist arguments from the Industrial Revolution

I've found Opus 4.8 and Fable 5 both difficult to learn from purely because of how annoying their writing style is. I'm finding GPT 5.6 Sol to be much better for this.

  • One nice thing but ChatGPT is that good image model means it can generate good infographics occasionally to illustrate. These become naturally compact in text.

I know this is not "as designed," but I kind of like it? Because, as long as it stays this way, it's still at least possible to tell if a human wrote something. Like, I know it's not much, but it gives you the ability to classify information as human generated or machine generated. Machine generated information may not be useless but it is different and IMO needs to be treated with a different level of skepticism. Not that you can simply trust human writing but it seems like the type of people who would publish machine writing have a different distribution of motives than those who would publish human writing.

I do think they're gonna figure out how to fix it at some point. :(

I think a signature Claude style of writing is good since it makes it that much harder to pass off Claude written text as human.

  • Easy enough to change. I have a Stylometry Skill fit to my preferred style—a mix of me and Terry Winograd. Give Opus 10 of your best paper thst you wrote and tell it to build a model of your style. hHard to distinguish except I make way more typos.

    • Let Opus write a tool that introduces statistically likely human typos. Don't let it just rewrite itself, let it write a model that it can apply.

      This sentence above filtered with the one-shot result of above prompt at a high typo-rate:

      Let Opus write tool that introduces statistically likely human typos. Don't let it just rewrote itself, let it wroite a model taht it can apply.

I found 4.6 more amenable than 4.8 to style directions, we'll see how 5.0 does. Super-small-sample-size: I think part of its "Claude-ism" style comes from its propensity to try and "proactively" move the conversation/work along. Not sure how this would fare in non-obviously-productive environments, I'd guess "it's still annoying" considering your evidence.

I'm also thinking of another benchmark: (quantified) stylistic range across different prompts. Just putting it out there if anyone wants to do the work for me :D

  • 4.6 is night and day better. It was before the big language switch up. Terrible direction that Anthropic has taken this.

I don't understand why Claude sounding like Claude is a bad thing?

What's next - complaining that `make` says "nothing to be done for 'all'"?

  • I'm very happy I can tell when AI wrote something. It keeps everyone honest. I'd be far more concerned if it didn't have a distinct tone and style.

I’m pretty sure Opus 5 is adapted to tricks from long reasoning in Kimi K3 and based on original Opus 4.8. It is not fable in any form.

  • Seems unlikely they adapted anything from K3 given the timeline of releases, similar to how K3 was obviously not distilled from fable