← Back to context

Comment by mrinterweb

8 days ago

Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.

Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.

  • I'm all for open models, but people seem to misunderstand what they are. They aren't the same thing as open source code!

    > open weights, open code and open data

    Even if you have all these things you still can't replicate a model because of randomness.

    You can backdoor a model with less than 1000 examples and it is impossible to detect.

  • Exactly what are the possible 'security issues' of self hosting an open weights model?

    • It may have been backdoored during training, potentially causing it to randomly start wreaking havoc at runtime, possibly in a clandestine manner (e.g. sneaking in bugs into generated code).

      4 replies →

  • > We need open weights, open code and open data.

    Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.

> could still hide backdoors that could be near impossible to detect.

But it won't change after you download it, so you can isolate those problematic cases and use another model for different use cases

In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.

  • It’s not hypothetical. Magic strings are a known and implemented feature for standard model interaction. Nearly impossible to detect unless you know where to look with current technology.

    • As long as I can say, "Model A, look for security holes in this code by Model B," I don't see this being a serious problem.

      It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.

    • Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.

      Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.

      11 replies →

which is moot point, if open model is hard to fully audit, then closed model is complete enigma and you should be more scared about closed models