Comment by _the_inflator

15 hours ago

Maybe check in the other direction as well: Astra to Fable.

I frequently simply let one of the three review what something that looks like awesome output by one AI gets totally annihilated by the other.

Finished outputs are easier to improve than bend a LLM to produce stuff like that in my observation.

Same with Gemini.

I yet have to find out how to handle this, whether I let agents check themselves and if on what process step.

Tweaking is hard.

I agree with your conclusion I am a huge ChatGPT and Codex fan, Gemini has to many infrequent quality changes when new models arrive ranging from great improvement to WTF.

ChatGPT seems to get scaling well while Claude still feels unstable, unclear usage statistics. Really weird.

Tough call I use all three.

I've had Claude review output from Qwen3.8-27b. Claude negged it! Said it was hallucinating!

27b may be small but it seems competent most of the time.