← Back to context

Comment by ricardobeat

16 hours ago

The main problem here is that a model that wildly fluctuates with 60% - 100% - 80% results will have the same wilson score as one that repeatedly scores 80% - 80% - 80%. So the 'confidence interval' bar is meaningless.

I'm not that well versed in statistics, but a standard box plot is probably the best alternative

A single result is binary. All we get from a run is which tasks were solved, which weren’t.