Comment by poytr1

3 days ago

I am thinking of where the initial training data came from. For example, Claude Code likely collected a substantial number of real-world coding trajectories through its CLI. However, trajectories involving tools such as LSP, MCP, or AST-grep were probably scarce in the dataset.

This lack of representation may also indirectly limit the effectiveness of subsequently generated synthetic data.