Comment by verdverm
1 month ago
They are more than opaque blobs, some examples:
- One can load them up in a model explorer to see the layers and other components, how it is designed
- One can fine tune the models, which requires adding LoRA to the model and then running some training iterations
You can make parallel arguments for binary executables.
model weights are more like a video file which has metadata about the format so any video player can play it
we run and change llm models with a variety of tools
Here's my thinking:
Cloud:
- You cannot directly execute a remotely-hosted program.
- You cannot run inference on an API-served model.
---
Closed-source:
- You can execute a program with the binary. You cannot generate a new binary, but you could try to reverse-engineer it or (painfully) modify its execution.
- You can run inference on a model with the weights. You cannot re-produce a new set of weights from scratch, but you can fine-tune.
---
Truly open:
- You can freely modify the source and produce new binaries.
- You can use the original training data and model architecture to independently re-produce the weights (assuming you've got the compute). You can modify the model architecture to get the weights that would've resulted from training the model that way.
---
To me these are pretty clear parallels... I don't think the weights provided in a vacuum are in the spirit of open source, historically speaking.
The policy argument is totally separate, of course, and I fully understand why none of the frontier labs are truly open.
2 replies →