And, yet it outperforms Qwen 3.6 35B A3B on Terminal Bench and SWE, etc. I dunno.
Edit: I guess you're right; apparently it's for simulation. I didn't look into it beyond the benchmarks. But, it does work in an agentic context, regardless. It'll write code, and drive an agent.
And, yet it outperforms Qwen 3.6 35B A3B on Terminal Bench and SWE, etc. I dunno.
Edit: I guess you're right; apparently it's for simulation. I didn't look into it beyond the benchmarks. But, it does work in an agentic context, regardless. It'll write code, and drive an agent.