PyPI Blog: Releases now reject new files after 14 days

3 days ago (blog.pypi.org)

The remaining risk now is that a patient, malicious actor could put out a new, clean source-only release, wait for ~7 days for people to decide it's safe and update to that version (and pass typical update delay controls), and then attach a bunch of malicious binary wheels. 14 days still seems to be too long.

Of course, this is already miles better than the current state of affairs where an old but popular package could become an infection vector at any time.

I’m a bit surprised this is possible in the first place. I get that you might not be able to upload everything in one go, but it feels like you should “start” and “finish” a release in that case, and once it’s finished you can’t modify it.

I guess the use case is that you might want to build a wheel for an older release for a newer version of Python?

  • The use case of uploading new wheels to an old release was actually (kind of) an accident: PyPI’s current upload API is stateless and originally there was only one file per release (sdists), so there was no need for a start-finish transition for releases.

    (This will hopefully change pretty soon, with the “upload 2.0” work.)

  • IMO a release workflow would be better with something like this:

    1) Upload all files in a staging state. Can be done asynchronously via multiple build hosts. Files in this state are referenced via their cryptographic checksums (e.g. SHA, Blake, etc)

    2) Make visible with a single call by providing a manifest with checksums for all source dists/wheels contained in the release. All artifacts are made visible atomically and the release is immutable.

    That allows authors to prepare uploads over however many days they need to coordinate hardware, but doesn’t allow for users to discover a release in a partial state.

14 days is still too long if you ask me. Releases should be immutable.

  • Published files within a release are immutable.

    The time limit is needed because a release can contain different binary wheels for different architectures.

    Consider the simplest case: your releases go out via GitHub Actions and separate wheels are built on the Windows, Linux, and macOS runners.

    Those won't all end at exactly the same time, so you need a release window during which they can finish and upload their generated files.

    That window used to be unlimited, now it's 14 days.

    That might seem like a long time, but it means more manual release processes still have time to coordinate, or release processes that need access to less common hardware that might require queuing for a while.

There seems to be a severe lack of hash pinning in "modern" software ecosystems. We figured out how to do it 20+ years ago with Git bringing hash addressed storage to the masses. Coming from a different background it was very surprising for me to see things like docker images, packages and github actions being updated at the whim of upstream registry. I much prefer the philosophy where builds are fully offline and predictable, even if not fully reproducible.

  • I’m not sure what this has to do with TFA: Python does have hash-pinning. TFA is not about modifying existing files on the index (PyPI doesn’t allow that), but about adding new files to a pre-existing release. But that doesn’t change the hash of older distributions on that release.

    • 1. There are multiple levels at which things are hashed, signed and checked when it comes to Python packages. Eg. each file in the Wheel beside the RECORD file is hashed using SHA256. This hash is never checked :( You can also sign your packages (RECORD.jws anyone? Is that still supported?), but nobody checks that either.

      2. There are hashes in the HTML served by PyPI. These are updated at the whim of both the index and the publisher. Even though they are checked by pip during install, they are worthless.

      3. There are many ways to install packages that work around (2). Custom index server doesn't have to provide hashes, and pip will happily install that. You can install from sources, from a package you've downloaded somewhere, form VCS, you can build it during install, all without even prompting the user to confirm the very scary choices.

      NB. I have no idea how do you make the leap from "adding files to release" to "not modifying the release". To me, adding file to release is sure as hell modifying it. Here's a very simple malicious example:

      I release package "innocent" with an empty "scripts" section. Then, in the subsequent modification to this release, I add the "scripts" section with a script named "notebook". Now, whenever my user wants to run Jupyter notebook, they will call my "notebook" program, not the one from Jupyter package.

    • I feel you’re quibbling over semantics here.

      In concept why can’t the full set of files in a release be a single, one-way hash value, with both adding or releasing changing the hash value?

      4 replies →

  • Would not count that as 20 years of sticking to that philosophy, though.

    We also figured out 20 years ago that SHA1 was not quite as strong as initially estimated, and not quite 10 years ago that generating two colliding documents was merely a matter of some serious computing power. A few projects went ahead and changed the name of their master branch, but SHA256 preference remains elusive.

  • Hash pinning (already) works, and this change is all about when you as a PyPI user do not use hash pinning for installing releases, when you pin just release version for example.

    The release consists of one sdist and zero or more wheels. Until now you were able to upload additional wheels at later time.

    • Very pedantic of me, but I figure it’s interesting to note: technically a release on PyPI can have zero files or even one or more wheels but no sdist. The former is a degenerate case that users don’t normally see, and the latter happens if the user chooses to only upload wheels (or their sdist upload fails for whatever reason).

      (This doesn’t change your observations at all! Just as a demonstration of how Python packaging’s data model can be unintuitive.)

      1 reply →

> To quantify how disruptive this change would be to existing workflows, the PyPI database was queried for projects that have published new files to old releases

While this may quantify how disruptive the change would be to those projects that are able to and do upload additional binaries to PyPI later, it fails to quantify how many projects already completely circumvent this block before it is even introduced.

e.g. If you tell pip to install from source.. the result may already be that you install a binary that PyPI never saw. A common hack for dealing with NVidia internals, which can explode into a large CUDA major version x GPU arch x platform x implementation x python_version cartesian product. The "extras" mechanism is not quite sufficient to model such combinations.

sample code: https://github.com/Dao-AILab/causal-conv1d/blob/4f6ae4e26ae5... https://pypi.org/project/causal-conv1d/

This seems like common sense configuration management 101. If I download v1.2 and it’s been published then it should be considered released and not modifiable. With exceptions for ‘dev’ releases of course. I have never published anything on PyPI but I would expect there is a publish button and finalize (?) optional button that if not checked after 14 days makes it final ?

  • Are there any package managers that have that kind of publish/finalize flow?

    Every one I’m aware of works either as a one-shot (you have to submit everything in one push) or lets you keep adding new assets forever (other, obviously, than PyPI with the addition of this 14 day wall).

Python packaging, the most convoluted way of creating simple zip files imaginable.

The tool fragmentation is insane, the demand to create "source distributions" was maybe funny in 2002 but just a hindrance now.

Packages no longer build since distutils was ripped out and upstream replaced it with meson etc.

Since building from source no longer works, which is profitable for third party vendors like Conda, "wheels" are uploaded. And they cannot be built on the server since the whole "scientific" ecosystem is perpetually broken. And they are separate artifacts, leading to the above problem.

Shipping checksummed tar archives is of course it not possible, that would hurt the income streams of the package profiteers.

The other consideration that would be useful is an explicit api for a developer to freeze the release, to prevent new file upload.

Kinda curious why releases just aren't fully immutable? Sane semver would dictate any update should at least be a new patch release.

  • The files in a release are immutable, but a release on pypi consists of multiple files for binaries that is a cross of architecture, os, and python version. Per other comments the upload api is stateless. The consideration is an attacker adding new files to an old release

    • GP wasn't asking about the files being immutable, whatever a mutable file would mean. They were asking about the releases. I'm curious too.

      The way it's done with rubygems, if you messed something up with the gem (pushed secrets, etc.), you "yank" (remove) the release and push a new one (different version). You can't work with the files in a release once it's been pushed.

14 days still seems way too long to me. As a user I thought releases on pypi were immutable!

  • How it works in practice is that some release flows add wheels for different platforms, as they get ready, separately.

  • Files are immutable on PyPI, releases are not (because releases are comprised of a set of files, and files are uploaded one by one).

    This is unintuitive, but the TL;DR is that files will never change on PyPI, but (previously) a user could upload a new file to a release years after their last upload to that release. This has some legitimate use cases (like allowing people to support new Python versions without bumping a package’s version), but also makes introduces challenges around locking and release security that are elaborated in the thread linked by the blog post.

    I agree this could probably be ratcheted down from 14 days over time, though.

    • Worth noting, because I think it’s confusing some folks on this comments page, that “file” here is “a wheel for a platform” or “a complete sdist”, not “each individual .py file in a release”

      You can’t go back later and add “evil.py” to a bunch of existing release files, but you could previously go find a bunch of releases that didn’t have arm64 files, publish malicious ones, and use that to catch people using those versions on arm64 systems

    • > (like allowing people to support new Python versions without bumping a package’s version)

      ...why Python is just breaking compatibility so bad with new version it needs that ?

It's kind of hilarious how everything that has to do with Python is so obviously wrong, with the obviously right way of doing things being right there, on the service, having been around since before Python even existed... and yet, Python will fail spectacularly every time.

If you ever used Maven, NPM or... I can't think about any other tool that doesn't automatically check checksums and signatures. Any Linux package manager ever used... Python's Wheel format has provisions for checksums and signatures! But they aren't checked.

Instead Python gets absurdly ineffective workarounds that will probably inconvenience a few developers and will do zilch for users.