Comment by lumost

2 years ago

The paper is interesting, but it seems to focus on iteratively training models on synthetic copies of the same data. Obviously, this is going to cause problems.

They did not address what happens if the model is trained on synthetic data that is distinct from the source corpus.