← Back to context

Comment by NitpickLawyer

3 days ago

> Why is Fable not on here?

Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

How do they handle these assurances? Personally I have zero trust in the AI companies not trying to use this data to get ahead in the game, and short of sharing the weights and harness so that the benchmarkers can run the models themselves, I don't see a satisfactory solution with this mindset.

  • OpenAI's Zero Data Retention claim held up in court. They were unable to produce prompts and outputs because they were never retained.

    I believe that is only available through Enterprise API for both Anthropic and OpenAI.

Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.

  • I think you could have accessed Opus on AWS then u don’t have to trust that the data will go to Anthropic?

    Just like the hugging face incident, Opus 5 could have escaped and went to grab data for training it shouldn’t have been able to..