Yep. An entity in the "dirty room" reads the thing to be reimplemented and produces a document that thoroughly describes its behavior. That document is passed to the entity in the "clean room" whose only knowledge of that system is through that document. Reverse engineering is legal, plagiarism is not.
Despite the fact that the raw output of the system is incomprehensible to humans, scanning a photograph of Mickey Mouse and running it through a lossy compression system like JPEG doesn't suddenly make it not a picture of Mickey Mouse. Similarly, running the code for a system through the lossy compression system that is LLM "training" doesn't suddenly obliterate that data and make that LLM a "clean room". If one has any doubt of that, remember that they are known to reproduce their inputs, even after all these years of work to make them not do that. [0]
With clean room you just have access to the api specifications, not the internals.
Yep. An entity in the "dirty room" reads the thing to be reimplemented and produces a document that thoroughly describes its behavior. That document is passed to the entity in the "clean room" whose only knowledge of that system is through that document. Reverse engineering is legal, plagiarism is not.
Despite the fact that the raw output of the system is incomprehensible to humans, scanning a photograph of Mickey Mouse and running it through a lossy compression system like JPEG doesn't suddenly make it not a picture of Mickey Mouse. Similarly, running the code for a system through the lossy compression system that is LLM "training" doesn't suddenly obliterate that data and make that LLM a "clean room". If one has any doubt of that, remember that they are known to reproduce their inputs, even after all these years of work to make them not do that. [0]
[0] <https://news.ycombinator.com/item?id=49727685>