Comment by fph
3 hours ago
To be fair, you picked a well-known tricky benchmark for LLMs: When working on an embedding spelling disappears after the embedding level. I imagine modern frontier models have tools that let them read back their input to work around this issue.
No comments yet
Contribute on Hacker News ↗