Comment by wg0
6 hours ago
Can I put it as Air Traffic Controller? With similar error rates as humans?
That would be the litmus test.
"Does not hallucinate" is not the same as "is never wrong".
So the ATC test could be the benchmark.
6 hours ago
Can I put it as Air Traffic Controller? With similar error rates as humans?
That would be the litmus test.
"Does not hallucinate" is not the same as "is never wrong".
So the ATC test could be the benchmark.
Not hallucinating is easy when you don't produce strings.
Hallucinating as we use the word really only applies to generative AI. Non generative AIs can't hallucinate, they can just be wrong.