Comment by walrus01
2 hours ago
Keep in mind that lm studio is just a GUI with llama.cpp/llama-server under the hood, so if you want to compile the latest llama-server and load your own choice of model, then connect to it with pi, opencode, etc or your own other choices of harness, that's also a popular option.
Things are moving fast enough these days that llama-server needs to be built from source every 4 or 5 days to keep up with model support and various tweaks in published quantized GGUF files.
Additionally there are a few different tweaks/branches of llama.cpp/llama-server that you can grab and compile to take advantage of changes people have made specific to discrete models and/or types of GPUs.
No comments yet
Contribute on Hacker News ↗