Comment by embedding-shape
9 hours ago
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
9 hours ago
This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.
The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.
Isn't this basically the mythical man-months LLM edition?