← Back to context

Comment by perching_aix

3 hours ago

What on earth do you even do with these models?

Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions?

I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you.

Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops.

I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really?

I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code.

  • I equate most LLM work to squeezing a glue bottle

    It's just glue code

    It's not complicated. Someone just has to be there to squeeze the bottle

  • It's possible it's my role distorting my perception, cause technically I don't write software, I work an SRE role. None of my items come pre-chewed or paced, it's all good luck and god bless.

    I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic.

    I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia.

    It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers.

    • For the last few weeks I've been using Luna with Medium Reasoning for routine debugging, installing/updating dependencies, test creation/fixes, CLI/config miscellaneous problem solving, etc and it's been solid. I can let it churn away for a half hour and it barely moves the needle on remaining usage on my Plus plan. Previously I had been using Sol Low/Medium and would frequently hit the 5 hour limit doing those kind of tedious automation and routine problem solving tasks.

    • Our ai basic analysis for SRE / k8s based platform is haiku and its surprisngly good.

      I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me.

      When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now.

      1 reply →