← Back to context

Comment by jchw

11 hours ago

I'd personally like to know more about what tools it used/wanted and the harness setup, because this sounds pretty cool. I have a dual Arc Pro B70 setup and currently get around 22 t/s which isn't great but isn't terrible either (it is at least less quantized.)

I've seen GPT 5.6 Sol happily invoke objdump and even write jobs to run headlessly which Ghidra when trying to disassemble a binary.

My M5 Pro gets around 12-15 (6 bit MTP), although I haven’t worked on optimising it at all yet.

A nice thing about running locally is you can run an uncensored model and you don’t have to worry about TOS violations on your OpenAI account when you ask it to “reverse engineer this ancient router firmware and give me a licence key that will work on it”.

  • Qwen is very much censored. Just try asking it about Tiananmen or how to build a bomb. But it is nice that you can experiment with it locally without having to worry about your account getting nuked

    • You are misunderstanding what they said, they are saying you can use uncensored variants of models like Qwen when running locally. There are quite a lot of people working to "uncensor" open weights releases. It seems to work although it would be nice if some third party was benchmarking the uncensored variants regularly to give us an idea of how well retained their skills are.

I added a line to address this, sorry it wasn't there before! It was Pi and only used Bash-based tools.

  • Cool. I was thinking of running Qwen3.8 through Codex, but maybe it's time I take a look at Pi.