Comment by oli5679
4 hours ago
You need a set of Evals, which catch enough of the mistakes your llm makes following your baseline prompt, so you have an equal or lower error rate than humans.
This can be determined by offline benchmarks if you build a system that takes small sequences of actions, and requires live ab test for long sequences of actions.
The more your human ops team work from documented standard operating procedure, rather than tacit knowledge, the less you need Evals except to capture edge cases
No comments yet
Contribute on Hacker News ↗