Comment by bitexploder
21 hours ago
Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.
21 hours ago
Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.
Right, imagine if instead they had coined new terminology that was not obvious and it re coined that - this would be close to a smoking gun
Afaict that didn't happen so there's just lots of speculation
One off might not work but how many n off you have to be is probably smaller than you'd guess, because the model does need to fit cases that are rare and would not be represented well in training e.g. esoteric things or very recently documented things.
You can probably game the metrics that models use to weight potential knowledge akin to SEO. Maybe have some bots parrot your data around a bit in some places online, maybe the model picks up on this and sees it as high engagement and promotes it over the correct data.
Maybe there are ways you can coax out the most optimal way to break into the training set out of the model itself.
Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?
That is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.
"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.
1 reply →