Comment by aroman
7 days ago
You have hundreds of hours with a model that was barely even released hundreds of hours ago?
The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
You could run hundreds of agents in parallel and arguably that counts.
No, it wouldn’t. The hours in question are human experience, not that of the agent.
Why not? You have far more results to review.
If you run a model on slower hardware are you getting more experience? Surely its a factor of model output reviewed and not human time.
1 reply →