← Back to context

Comment by ronbenton

1 day ago

> 600,000 ServiceNow tickets are another element of the collection, along with 13.7 million active emails addresses from Oracle’s Responsys marketing application, and details of 11 million sales of in-flight Wi-Fi services.

I really doubt all this stuff was “de-identified”

I don't see how it's possible any more, when correlated against all the various other data sources. And a record that might be unidentifiable now might become unique with more correlated data sources.

  • All you need is gender, birthday, and zip code to uniquely identify ~85% of the Americans (per some study). There is far more data in these records, even before AI there were far more detailed marketing profiles of people, and with AI software, it's likely easy to use it to identify people at scale.

Most laws are designed to accomplish specific objectives. For example, there might be laws about whether you can collect certain kinds of data, for the sake of privacy or protecting some attribute. As technology improves, however, it may become possible to collect different kinds of data (that are not regulated to not be collected) and still recover the original attribute.

For example, 20 years ago, few people would have imagined that you could build a fingerprint of a person from all their behaviors on the internet. You didn't need to regulate away certain kinds of data collection to prevent such fingerprinting because it wasn't possible. With ML and AI, it has become possible.

My point is that as technology improves, "de-identification" takes on new meaning. Whatever de-identification was 20 years ago is not what de-identification is today. And whatever de-identification was used here is probably insufficient to guarantee that customers can't be identified.

We need new regulations and we need them yesterday.

> de-identified

De-identified but far from useless.

as an example, they can remove the names off these sales data, so you can't identify who purchased what items. However, the purchaser would be identified by some sort of number, and you would be able to extract information about purchasing habits, and aggregate these habits into usable information for advertising purposes (like targeting and profiling).

And that's before AI training for LLM purposes.

Wello this is troubling. How much other data must they have bought that wasnt public