← Back to context

Comment by Theory42

9 hours ago

A useful critique, thanks.

I would say that the grader has the same threshold for whatever answer it receives, and is equally harsh on whichever it grades. Any scores above the baseline (1.0) are really claims about parity, rather than better understanding in the compressed format.

The decoder step is a model expanding the cablese to regular text, not having seen the initial question. A separate model instance then reads that regular text and answers. And the result is still at parity with the plaintext record.

Had cablese knocked out information, that wouldn't have been the result, would it?