Comment by shahidhussain

4 days ago

Author here. I bought this CD in 2011 and spent 15 years failing to find out who made it. The band charted at #47 on the UK Independent Singles Chart in June 2006, so it was a real, professionally made, commercially distributed record, and I couldn't find anything online: no credits, no fingerprint match, no company registration for the label, nothing on streaming.

Part of the answer turned out to be that their website was Flash, so it was never meaningfully indexed, and their MySpace was never mirrored. A band can be entirely findable in 2006 and entirely gone by 2026 depending on which technologies they happened to pick.

It was eventually solved by searching on the producer's credits rather than the band's name, and then by two people replying to cold emails. Both the producer and the singer were generous enough to be interviewed, so there's some fresh detail in there about the recording session as well.

That's an epic search but at least of all a wind at the end.

Sadly the ability to search via the usual suspects has diminished greatly from 2015 onwards, not all of it is on lack lustre efforts of search engines involved, some of the data had become just too hard to get from within forums where human intelligence rules. Of course the lack of competition from new search engines isn't good either, too many sites aim to block all but the top handful of bots from their content ...

  • Thanks for the note. Yeah, it was almost impossible to find the data without a ton of help. I’m sure things will change again as bot internet usage overwhelms human usage — maybe more silos, maybe more bot only interfaces. Interesting times!

I am one of the developers behind the Official Charts site, so thank you for posting the story. It's heartwarming to know that something I have worked on has helped you on this journey - even if it's in a very small capacity, maybe you wouldn't have connected the dots otherwise. Hopefully this means my work has helped others in the same way!

> A band can be entirely findable in 2006 and entirely gone by 2026 depending on which technologies they happened to pick.

This is interesting, and in some cases desirable. I wonder what technology choices one could make today if they wanted to have an internet presence now but not have it stored forever. Would that even be possible today?

  • You can block the Internet Archive with a robots.txt file, but no guarantees that other archives/scrapers or even individuals won't index or copy it. But for AI training, what exists in training data probably isn't the full site, and unlikely that an LLM could be spit back out verbatim (that's a guess).

A good read, thank you. There are many excellent bands and songs gone through time. It can be difficult to discover rare songs, or even slightly out there ones, through streaming apps like Spotify.