Comment by senko

4 days ago

>> The models are nondeterministic, and therefore it's pretty normal for different runs to give different results.

> And how is that an excuse? […] this qualifies as strong evidence…

This qualifies as nothing due to how random processes work, that’s what the gp is saying. The numbers are not reliable if it’s just one run.

If this is counter-intuitive, a refresher on basic statistics and probability theory may be in order.

1 comment

senko

bsder 3 days ago

> If this is counter-intuitive, a refresher on basic statistics and probability theory may be in order.

I'm not running "statistics". I'm running an individual run. I care about the individual quality of my run and not the general quality of the "aggregate".

The problem here is that the difference may not be immediately observable. Sure, if it doesn't give a correct answer, that's quickly catchable. If it costs me 10x the time, that's not immediately catchable but no less problematic.