... the API is geared at requesting specific content items (posts, comments, users). There doesn't seem to be a way to directly make a request for a front-page history page (that is, the 30 items archived on a given date. Say, 2008-11-05:
I could look more into their methodology to see if I can use similar approaches.
The existence of "dead" and "deleted" values does seem interesting. I might do some playing with those to see what shows up (I suspect that most additional information is suppressed...)
That's missing the title and URL, as I suspected it would, though the submitter UID is available.
To get top stories by date I'd actually have to submit more requests, walking through item numbers, splitting out comments and stories. Based on Whaly's 2021 retrospective, with about 4.2 million items (stories + comments) posted in total, that's about 12,000 items per day. Versus, well, one "Past" page result...
I know that.
I've not worked with the API, and there's the blessing/curse (blurse‽) that HTML is a known, if poor, standard.
API always translates to "one more thing to learn, that's applicable to a single-use case". HTML scraping / sorting I can apply across multiple sites.
That said, a standard, say, JSON packaging of website contents available on request might be fun to have.
I feel less bad hammering firebase in a "while True:" loop vs hitting HN's servers.
12,000 times less bad? <https://news.ycombinator.com/item?id=48868910>
3 replies →
Looking at the API ...
... it's starting to make sense, but ...
... the API is geared at requesting specific content items (posts, comments, users). There doesn't seem to be a way to directly make a request for a front-page history page (that is, the 30 items archived on a given date. Say, 2008-11-05:
<https://news.ycombinator.com/item?id=29769470>
I could look more into their methodology to see if I can use similar approaches.
The existence of "dead" and "deleted" values does seem interesting. I might do some playing with those to see what shows up (I suspect that most additional information is suppressed...)
OK, looking at a recent dead atomic128 comment:
So userID is visible.
And from a current dead submission in the New queue:
That's missing the title and URL, as I suspected it would, though the submitter UID is available.
To get top stories by date I'd actually have to submit more requests, walking through item numbers, splitting out comments and stories. Based on Whaly's 2021 retrospective, with about 4.2 million items (stories + comments) posted in total, that's about 12,000 items per day. Versus, well, one "Past" page result...