← Back to context

Comment by 0x457

10 hours ago

Some classifiers are tiny - like 2k params, maybe even less.

This whole thing started because I wanted something to help me play Dune Imperium. Even relatively large models with vision encoders couldn't reliably extract the full state of the board. Now that I have ~2k labeled screenshots, I want to train heads on top of SigLIP2 to extract all of that data in one go.

That's how it started. Now the thing supports multiple kinds of datasets:

  Images - currently the Dune Imperium and Bolatro screenshots, with SigLIP2 heads being the next step.

  STT - my self-hosted Linux dictation tool feeds this dataset. I run Nemotron ASR tuned for my voice.

  TTS - for Piper TTS, trained to speak like SHODAN. Trained from data generated by Qwen3-tts + original video games files.

  Text pairs - for a 1.2B model that converts normal text into "what would SHODAN say?"


  FastApply - a Qwen3.5-4B LoRA adapter for doing fast edits.

  Chat threads - all agent/chat threads get saved too, so eventually I can turn the useful ones into a dataset and train a LoRA for a really good Rust-specialized version of Qwen3.8-27B.

  Tool calls (extracted from chat threads) - this is where I want something Jev-like, mainly to add an auto-approval mode to my agent harness.

A model router isn't planned because I'm trying to gear everything toward self-hosting, and there just isn't that much to route between. I’ll probably build something Jev-like for smart-home control, though.

The FastApply dataset is already ~20k entries, with the majority of outputs being 8k–16k tokens. The STT dataset is roughly 30 hours and growing.

Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need.

> Basically, the whole thing has turned into a Collect -> Distill -> Train pipeline for whatever I happen to need

Amazing, thank you for sharing your setup. Very cool applications