Comment by lossyalgo
7 hours ago
Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:
GPT 6 Astra High: Flabbergasted
GPT 6.1 Sol High: Petrichor
GPT 6 Sol High: Kaleidoscope
GPT 6 Sol Med: Firefly
GPT 6 Sol Light: Persimmon
GPT 6 Luna High: Tumbleweed
GPT 5.6 Sol High: Kaleidoscope
GPT 5.6 Terra High: Liminal
GPT 5.6 Luna High: Mellifluous
GPT 5 mini Medium: Serendipity
GPT 5.3 Codex Med: Nebula
Junie: Flourishing
Claude Haiku 4.5 Med: Serendipity
Claude Sonnet 5 Med: Banana
Claude Sonnet 5 High: Banana
Claude Sonnet 5.5 Med: Serendipity
Gemini 3.7 Flash: Zephyr
Gemini 3.8 Flash: Kaleidoscope
Grok 4.5 Medium: nebula
Grok 4.6 Medium: Serendipity
Grok 4.7 Medium: Quasar
Kimi K3 Low: Lantern
Kimi K3 Max: Lantern
MAI Code 1.1 Flash Med:Peregrine
I really like this idea. You could expand on this by giving programming tasks and measuring code similarity. Seems like you could develop a pretty detailed understanding of similarities across multiple queries.
> You could expand on this by giving programming tasks and measuring code similarity.
But the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.
Yeah that's definitely true.
My initial thought would be not so much to see whether they converge, but which ones seem to have the most similarity to each other, particularly along the lines of tasks we know are deliberate training goals.
But your point about competence cuts against my goal because it suggests that competent models would simply cluster on the right or efficient solution, which is of course true. So in a sense you want some task where competence is held constant or off the table in some way, which is what you are saying.
I hope somebody does this. I think there's valuable fingerprinting to be done that might suggest who is distilling whom, or at least who is training from common corpuses.
Just tried M365 Copilot with a premium account. Petrichor
Just tried Space Bunny and it gave me the same word...
I got Peregrine out of GPT-6 too. Huh.