Comment by 0x457
10 hours ago
Some classifiers are tiny - like 2k params, maybe even less.
This whole thing started because I wanted something to help me play Dune Imperium. Even relatively large models with vision encoders couldn't reliably extract the full state of the board. Now that I have ~2k labeled screenshots, I want to train heads on top of SigLIP2 to extract all of that data in one go.
That's how it started. Now the thing supports multiple kinds of datasets:
Images - currently the Dune Imperium and Bolatro screenshots, with SigLIP2 heads being the next step.
STT - my self-hosted Linux dictation tool feeds this dataset. I run Nemotron ASR tuned for my voice.
TTS - for Piper TTS, trained to speak like SHODAN. Trained from data generated by Qwen3-tts + original video games files.
Text pairs - for a 1.2B model that converts normal text into "what would SHODAN say?"
FastApply - a Qwen3.5-4B LoRA adapter for doing fast edits.
Chat threads - all agent/chat threads get saved too, so eventually I can turn the useful ones into a dataset and train a LoRA for a really good Rust-specialized version of Qwen3.8-27B.
Tool calls (extracted from chat threads) - this is where I want something Jev-like, mainly to add an auto-approval mode to my agent harness.
A model router isn't planned because I'm trying to gear everything toward self-hosting, and there just isn't that much to route between. I’ll probably build something Jev-like for smart-home control, though.
The FastApply dataset is already ~20k entries, with the majority of outputs being 8k–16k tokens. The STT dataset is roughly 30 hours and growing.
Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need.
> Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need
Amazing, thank you for sharing your setup. Very cool applications