Comment by timr

10 hours ago

> We don't have enough accurate knowledge to say that, and it doesn't seem to be the case at all.

The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".

Only if you interpret statements as being binary logic.

"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.

Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"

  • Yes, you were guessing. That's the only thing you could be doing, since, as you said, nobody actually knows.

  • We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:

    > We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].

    But it is in no way equivalent to this:

    > We can't say that for sure, but my money is on it not being a simple case of intellectual property theft

    That is a different sentence.

  • > not being a simple case of intellectual property theft

    No, it's an aggravated case, since it's the same way they got all of their training data in the first place.

    • Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).

      Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.