Comment by walrus01
5 days ago
This is a concern but given its size, it's also going to cost a potential user $500-600k in hardware to self host and run Kimi K3 at any useful speed with full context size. It's not something that just anyone interested in attacking a system can use.
The size/cost of hardware is far beyond even something like a self-hosted GLM5.2 Q8 at approx. 850GB GGUF file on disk size, which can run at a slow tok/s rate on a server with 1536GB RAM.
Scam center operators in Myanmar have resources well over $600k, they are building entire office blocks for forced labor scam centers.
If the ROI is there I absolutely see them venturing into that direction.
Why would you caculate 500k?
if Kimi is around 1-3tb big, even current DDR5 prices are at 15k.
It remains to be seen once it's released, let's say theoretically unsloth quantizises it to their own version of Q8-XL, and it's 2TB in size. But we don't know what speed it will run on a dual or quad socket xeon server with, let's say 48 * 64GB DIMMs, 3TB of RAM. Enough room for the model and its full default context size. 10 tokens/s? What kind of speed will it run at when context fill is 200,000+?
The ability to run it fast enough to go on a recursive nested attack of finding an entry point into something and then proceeding with lateral movement/privilege escalation and such will require more speed, like 40-50 tok/s at least, unless you're prepared to wait weeks.
Same that some people are right now running GLM5.2 in its 850GB version on CPU-only and a pile of DDR4 or DDR5 server RAM, yeah it runs, but not very fast. Good enough to give it "build this piece of something and wait a few hours" tasks, come back later and see what it's done. Yeah, you can do that under $20-30k for sure. Even with something like a used Dell R940 with 1536GB RAM bought on eBay.
Because you need GPUs to run it fast enough for an attack to be effective. 8 GPU servers with enough VRAM are that expensive.
FWIW, you can get a 16x RTX 6000 Pro setup running for significantly less than that. Even considering the electrical hookup fees (it’s a lot of power and cooling). Home ownership might push you into the original quote though.
3 replies →
To run it "at useful speed"
Why would they need to selfhost? They can run it via API ala openrouter or runware or whatever.
Or on runpod with rented gpus or inference.
Outsourced LLMs are usually censored. It's presumed if you want uncensored, you have to run it yourself.
Jailbreaks exist and work...