Comment by workbreak
7 hours ago
Every tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads. This really bites when using expensive models since most models are 1/10 for cached input.
No comments yet
Contribute on Hacker News ↗