Comment by subygan
4 days ago
This does not really work well, if you don't know the complexity of the problem ahead of time and ensure all future conversations go to the same model.
Else, you break the cache by doing a round robin of the same conversation across different models. Likely you'll end up paying more than what it would've cost with a cache aware system
You can use the Ralph Wiggum technique: https://ghuntley.com/ralph/
Cache hits should be part of the strategy.