← Back to context

Comment by jefftk

1 day ago

I think you might have missed the deidentification piece?

Not trying to be snarky, and perhaps it wasn't well stated, but the last paragraph I said I'm concerned about identification via my writing style. If they have my emails, they would have my writing style. It doesn't have to be tied to PII there, they can cross reference it with my blog. I'm speculating because I read that you can identify people by a few sentences of their writing.

"Deidentification" seems really murky and imprecise at best.

  • Reidentification via writing style is definitely possible, and I doubt the vendor will modify things in a way sufficient to handle that.

    But I think this is a place where we should apply bounded distrust: there are lots of places where we should distrust Google, but reidentifying people in an explicitly deidentified dataset isn't one of them.

    • The re-identification is done by ML, at which point it's basically undetectable. If you give all this day to, say, and LLM, and the LLM also has your blog, the identification is embedded in the weights.

      I can ask "Tell me about Person A's experience with airlines" and it will tell me. Or I can ask more generally, "knowing your knowledge, derive a no-fly list", and then it's likely Person A will be on it.

  • >murky

    Maybe, a little. Top comment on an old HN stylometry tool:

      “Wow. This gives a lot of false positives, but it found all ~10 of my old accounts over the years.”
    

    https://antirez.com/hnstyle ). Anyway imagine Google’s funding + data + new techniques—shouldn’t be terribly murky these days.

You have to trust that this really "deidentifies". Time and time again it was shown, that the measures taken were not enough to anonymize.

E.g. the parent wrote that he fears, he could be identified by his writing style, which is totally plausible. How would you "deidentify" this?

  • Even if they follow to the letter a deidentification process, Google and Meta have so much data about individuals that re-identification shouldn't be very hard for the majority of airline passengers' data they put their hands on.

    Of course, takes a lot more effort than not doing proper deindetification in the first place but if they wanted to appear like caring about data privacy they still have enough data points to correlate the sets later on (and/or over time).

    • The idea that there’s a nefarious plot to do something super evil with this data is a bit crackpot though based on their incentives.

      Remember, Google = Ads. Their only focus and only care. Their mission statement, rendered accurately, is “Ads ads ads ads. Effective ads. Ads worth paying a lot for. Ads ads ads. Advertising and ads.”

      If they choose to be evil in some additional way, (1) remember, they would only do that if in some way it serves their advertising needs — not to offer innovative new black-hat databroker services to airlines, and (2) this little dataset will not need to be re-identified. They’ll just use the 20 years of email and search data they already have on like half the world’s population.

      1 reply →

Even before LLMs there were multiple papers written about ways to to reidentify people with ML and other statistical analysis. It is probably now even more trivial especially if you are Google.

I have a bridge for sale, hardly seen use, pay me ${money} and you can collect it in New York City. Interested?