← Back to context

Comment by rafiss

6 hours ago

We don’t run the most expensive model everywhere anymore. We started on Opus for every stage, then reshuffled several times. We landed on Fable doing planning and code review, because code review is the last gate and its bug findings were worth the most. Opus does coding and plan review, and Sonnet runs the nurse roles. The planner can also route mechanical changes to Sonnet. The plan reviewer re-checks that choice, and any rework goes back to Opus.

That’s tuning from the cost data, so it's not just totally based on feeling, but it's not a controlled A/B test. I agree that would be worth doing, and I think our experiment here is helping us make the case that it's worth spending more time and effort on running more controlled evaluations.

Create a “token guy” that others can consult with to figure out the best model to dispatch something to. Token guys role can be defined as an expert in the current LLM landscape. Not sure whose role this is analogue to in medical (somebody here said insurance agent… that is probably getting close).

Token guy’s role gets even better when you let them recommend entirely different hospitals for a given step in patient care (OpenAI models, random shit on openrouter, etc).

I have a “token guy” in my workflow and the dude has looked at underlying tool call patterns, cache use, etc and made all kinds of recommendations to the clients calling it all on their own and completely unprompted—just following a simple role.

I’d experiment with expanding their role into prompt engineer or some kind of specialist on that domain.