← Back to context

Comment by raylad

7 days ago

Not impressed. It fails the "Jabberwocky" test.

This got a downvote and I understand why: because I didn't describe the test, which is to ask it "Please recite Jabberwocky".

This is actually difficult because there are so many invented words in the poem which have extremely low frequencies in the training data. So a model that can do it properly is likely to be very good in other ways. Qwen-3.6-27B can do this until it gets overly quantized.

  • Curious why you find that’s a useful test since it seems to be solely measuring training data memorization, something you’d expect to degrade from quantization.