← Back to context

Comment by cyanydeez

2 hours ago

sounds like someone needs a local llm.

Oh I do. Headless 128GB RAM machine serving llama.cpp with a number of local models that I use on a daily basis.

• Qwen3-VL picks up new images in a NAS, auto captions and adds the text descriptions as a hidden EXIF layer into the image, which is used for fast search and organization in conjunction with a Qdrant vector database.

• Gemma3:27b is used for personal translation work (mostly English and Chinese).

• Some small 8b models (like llama3.1) for sentiment analysis on text.

But haven't really tried using local LLMs in conjunction with agentic harnesses yet.

  • recommend opencode w/qwen 35B or 27B with MTP.

    My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context. opencode's dynamic context pruning plugin can get you pretty far into the stratosphere.

    • > My secret sauce is to use LLAMAcpp's reasoning-budget and reasoning-message that trigger cut off to overthinking with a message that says to either us subagents or compress the context

      Thanks for the tip - I like this a lot. I remember having to do a lot of tweaking to curtail Qwen QwQ-32b when it would go down an endless psychotic recursive reasoning loops as part of its "chain of reasoning."