Comment by evilmathkid

5 days ago

well not if its an open textbook exam

in non metaphor terms: In many ML situations you can carry the train set with you test time. Eg: KNN, SVM, replay buffers, etc

this is one such case

--

the separate overfitting concern is fair, look at private holdout performance for that. it performs on par with TRM (a comparable model) in the private set, ofc with far less compute

I think you can do whatever you want with your train set, what would be concerning is training with the test set.

I don't think overfitting is separate, the consequence of putting the benchmarks in your training set is not that you "cheat" by breaking some moral code, it's that it breaks the purpose of the benchmark and trains your model to be good at that benchmark only, instead of being generally useful.