Comment by sarjann

7 hours ago

Sounds like this would be expensive on the input token side with some saving on the decode side. For workloads with large documents / context that could be an issue.

AFAIK they solve this through their RAG based knowledgebases to only search and use the most relevant Information.