Comment by plumb_samji
5 hours ago
Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
5 hours ago
Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness
No comments yet
Contribute on Hacker News ↗