← Back to context Comment by gitpusher42 9 hours ago Thank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots 2 comments gitpusher42 Reply WithinReason 7 hours ago That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger? gitpusher42 41 minutes ago Check for colibri, dwarf star and flash-moe. they do similar things with bigger modelshttps://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe
WithinReason 7 hours ago That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger? gitpusher42 41 minutes ago Check for colibri, dwarf star and flash-moe. they do similar things with bigger modelshttps://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe
gitpusher42 41 minutes ago Check for colibri, dwarf star and flash-moe. they do similar things with bigger modelshttps://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe
That's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?
Check for colibri, dwarf star and flash-moe. they do similar things with bigger models
https://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe