Take ModernBert, add a fully connected layer with 255 outputs, take a bunch of classification datasets from Huggingface, write the code to have the datasets fit the jev format on these 255 outputs, do supervised fine tuning on the datasets with that format. then use a confidence loss of some kind
Asking as a curious bystander: would that be sufficient?
I'm not familiar with ModernBert (my understanding stops around the original Bert), but it feels like this is asking it to do lots of heavy lifting. Can it do that much?
Take ModernBert, add a fully connected layer with 255 outputs, take a bunch of classification datasets from Huggingface, write the code to have the datasets fit the jev format on these 255 outputs, do supervised fine tuning on the datasets with that format. then use a confidence loss of some kind
Asking as a curious bystander: would that be sufficient?
I'm not familiar with ModernBert (my understanding stops around the original Bert), but it feels like this is asking it to do lots of heavy lifting. Can it do that much?
Well the numbers don't lie Laya is ModernBert-Large (421M) fine-tuned and has a similar to better performance than TypeSafe Jev https://raw.githubusercontent.com/NandhaKishorM/laya/main/as...
The real guide is here: https://huggingface.co/convaiinnovations/laya