Comment by stldev
16 hours ago
Can confirm. Max effort helps; limiting context <= ~20-25% is crucial anymore.
> * keep active sessions active. It seems like caches are expiring after ~5 minutes (especially during peak usage). When the caches expire it sees like all tokens need to be rebuilt this gets especially bad as token usage goes up.
Is this as opaque on their end as it sounds, or is there a way to check?
No comments yet
Contribute on Hacker News ↗