Comment by jmaker

2 days ago

This is an interesting hypothesis to test. Depending on the embedding model, such less frequently encountered character sequences in the coding domain could indeed produce less optimal tokens. A dedicated model trained from scratch on Haskell would be optimal but wouldn’t generalize well beyond that, but can serve as a deliberately overfitted baseline.

What I meant was really about the output tokens, on the API pricing level of interaction with an LLM.