← Back to context Comment by WithinReason 7 hours ago Nice job implementing expert caching! 2 comments WithinReason Reply gitpusher42 7 hours ago Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots WithinReason 5 hours ago That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?
gitpusher42 7 hours ago Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots WithinReason 5 hours ago That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?
WithinReason 5 hours ago That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?
Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots
That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?