Comment by Tallain
9 days ago
It's published but I would doxx myself by posting the link here -- I can email you if you'd like.
Also, full disclosure, because this was written entirely for myself, the code behind it is 100% LLM generated. I set the specs and direction but didn't touch code at all, just let agents churn while I did my day job and tested the results later.
For the recommendations:
- Seed selection is performed by building a mix of seed artists from recent listening, lifetime favorites, old obsessions, and underexplored artists.
- Seed weight combines `log(1 + play count)`, track diversity, number of listening months, and a recency decay
- Candidates are generated at the union of artists reached by "similar artist" edges through Last.fm and MusicBrainz-enriched local data (APIs are used to enrich DB initially, then after that it's all offline)
- Discovery excludes known artists, hidden and recently-reviewed artists
- Discovery requires edges with at least two distinct seeds, or one strong seed edge
- Discovery is scored by similarity (Last.fm + tag), feedback, and rediscovery thresholds. Then score is subtracted for excessive knownness within a genre, obviousness, recommendation fatigue, and shallow prior listening. The thresholds are by default but configurable:
```
0.35 × calibrated Last.fm similarity
+ 0.25 × calibrated tag similarity
+ 0.10 × rediscovery
+ 0.10 × feedback
```
- Serendipity is additional score provided when a candidate connects several seeds that are themselves weakly connected, or are in separate connected components. This is a weak implementation of the Auralist serendipity algorithm, and will be improved on.
- Finally, diversity reranking; a greey MMR-style reranker.
```
rerank score = λ × candidate score
− (1 − λ) × similarity to already selected results
```
Very cool, thanks for the write up on the recommendations. I'm always hungry for new music and I have a somewhat patchy last.fm history going back to 2005. I'd love to check out your tool, email in profile. Cheers!