← Back to context

Comment by hununu

4 days ago

Interesting. Have you repeated these experiments with recent models? I'm thinking frontier models APIs have tools/MCPs for math stuff but curious about recent Qwen models, etc.

Haven't tried recent frontier ones, Astra for example doesn't support logprobs emission on the API, and Sol and Luna supposedly support it with reasoning disabled. Haven't tried local models like Qwen 3.8 27b (I'm actually exploring their thinking trace, it's a lot of fun)