Comment by fathermarz

18 hours ago

I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.

We don’t run the most expensive model everywhere anymore. We started on Opus for every stage, then reshuffled several times. We landed on Fable doing planning and code review, because code review is the last gate and its bug findings were worth the most. Opus does coding and plan review, and Sonnet runs the nurse roles. The planner can also route mechanical changes to Sonnet. The plan reviewer re-checks that choice, and any rework goes back to Opus.

That’s tuning from the cost data, so it's not just totally based on feeling, but it's not a controlled A/B test. I agree that would be worth doing, and I think our experiment here is helping us make the case that it's worth spending more time and effort on running more controlled evaluations.

  • Create a “token guy” that others can consult with to figure out the best model to dispatch something to. Token guys role can be defined as an expert in the current LLM landscape. Not sure whose role this is analogue to in medical (somebody here said insurance agent… that is probably getting close).

    Token guy’s role gets even better when you let them recommend entirely different hospitals for a given step in patient care (OpenAI models, random shit on openrouter, etc).

    I have a “token guy” in my workflow and the dude has looked at underlying tool call patterns, cache use, etc and made all kinds of recommendations to the clients calling it all on their own and completely unprompted—just following a simple role.

    I’d experiment with expanding their role into prompt engineer or some kind of specialist on that domain.