Comment by dredmorbius

2 months ago

Looking at the API ...

... it's starting to make sense, but ...

... the API is geared at requesting specific content items (posts, comments, users). There doesn't seem to be a way to directly make a request for a front-page history page (that is, the 30 items archived on a given date. Say, 2008-11-05:

<https://news.ycombinator.com/item?id=29769470>

I could look more into their methodology to see if I can use similar approaches.

The existence of "dead" and "deleted" values does seem interesting. I might do some playing with those to see what shows up (I suspect that most additional information is suppressed...)

OK, looking at a recent dead atomic128 comment:

  $ curl -s 'https://hacker-news.firebaseio.com/v0/item/48820709.json?print=pretty'
  {
    "by" : "atomic128",
    "dead" : true,
    "id" : 48820709,
    "parent" : 48819517,
    "text" : "[flagged]",
    "time" : 1783444517,
    "type" : "comment"
  }

So userID is visible.

And from a current dead submission in the New queue:

  $ curl -s 'https://hacker-news.firebaseio.com/v0/item/48868688.json?print=pretty'
  {
    "by" : "millwright-sw",
    "dead" : true,
    "id" : 48868688,
    "score" : 1,
    "time" : 1783743361,
    "type" : "story"
  }

That's missing the title and URL, as I suspected it would, though the submitter UID is available.

To get top stories by date I'd actually have to submit more requests, walking through item numbers, splitting out comments and stories. Based on Whaly's 2021 retrospective, with about 4.2 million items (stories + comments) posted in total, that's about 12,000 items per day. Versus, well, one "Past" page result...