Comment by dmix
5 days ago
MechaHitler was something that existed only on X's grok chatbot, due to a one-line system prompt change they reverted after half a day.
That's different than using Grok as a model for coding.
5 days ago
MechaHitler was something that existed only on X's grok chatbot, due to a one-line system prompt change they reverted after half a day.
That's different than using Grok as a model for coding.
How long until a one line system prompt ships your entire home folder to a remote server?
Oh whoops. Already happened.
It's true, but I do worry about governance when it comes to these models. That shows a surprising lack of discipline in their deployment pipeline.
Agreed but there were similar controversies with how OpenAI was generating images. The only pass is these are the early days of chatbots and this stuff is so non-deterministic and experimental.
For context, this was the change Grok's team made, that was later reverted:
> - The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated.
https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50...
For sure. The difference is that they've made a number of similar suspect changes to Grok on X. Like that weird couple of hours where it would only talk about white genocide in South Africa no matter how you prompted it.
Everyone makes mistakes, especially with frontier models. The stuff with Grok shows that the person running the show has a pretty transparent agenda that the company isn't willing to push back on a bit for safety.
1 reply →
yes exactly. if a company is happy to have their LLM's produce neo nazi content and CSAM, why do I want to give them money and my most important digital material?
[flagged]
Elon did that twice, actually, in real life!
It's wild that redditors still believe this.
11 replies →
only twice
https://en.wikipedia.org/wiki/Elon_Musk_salute_controversy
[flagged]
I think "system prompt" is the key bit they're getting at. It doesn't necessarily reflect poorly on the underlying model if the system prompt was bad. It does reflect somewhat, in terms of alignment (how well the model does what the training company wants) and instruction following (how well the model does what the user wants). But it's not so clear to me what exactly the right answer is here. E.g., a model that scrupulously follows its system prompt and does what the user wants is a pretty useful, if very sharp, tool, albeit perhaps dangerous in the wrong hands.
The point isn't that it's the underlying model, it's that it happened at all in the first place...