Comment by dofm
2 days ago
Gemma 4 26B does really well at this specific task (and a general MySQL-related puzzle I test on). I rather like it and now they have fixed tool calling, I would use it. I think maybe it has been trained well with "consumer" programming languages like PHP that are sort of commonplace things people want to do. I think for less commonplace programming languages, maybe it's worse.
Muse Glimmer thinks well and codes well in my tests; it does fine at this. I really like it so far, but my tests are fairly shallow.
One thing I have been struck by — my prompt includes this sentence:
"Please read the following and then ask me any further clarifying questions you need before proceeding with code generation."
Almost all models I've tested interpret this as an instruction to ask questions regardless. Qwen 3.8 27B is the only one that either expresses confidence that it doesn't need to ask clarifying questions, or in higher reasoning effort ultimately asks questions, but offers up defaults I can choose with a simple reply.
No comments yet
Contribute on Hacker News ↗