Comment by weakfish
1 day ago
I’ve been trying to put my finger on what it is that is happening when the robot writes stuff like “3 campuses, one app” example, I’m glad the author was able to identify it as chat context leaking through.
The other version of it is the robot over-indexing on some part of the prompt and leaving comments places. An example is I ask it to prefer integration tests using TestContainers, it starts adding comments to every new test saying “Real services, no mocks” or something.
And yes, I have a line in my *.MD saying not to do that.
LLMs love https://tvtropes.org/pmwiki/pmwiki.php/Main/SuspiciouslySpec... . You tell them not to do a thing, they make a point of saying they didn't do the thing.
It's partly a "Don't think of a Pink Elephant" problem. I've worked on several projects where I had to keep telling prompt writers to stop writing negative examples because the more you include the more its "attention" to them is all it has. Like telling a toddler not to do something and being surprised that is now all they can think about and they want to keep doing it. These prompt writers kept getting surprised that I'd delete all their negative examples and harshly worded "Don't do X" and "Never Y" and "NO: Z" sections they spend so much time on and got better results with smaller more focused positive example only prompts.
Hah! I also thought about it this way and ended up added a "purple elephant rule" to my pi prompt to discourage the behaviour, since I figured LLMs lean on metaphors so much.
Of course, I quickly reverted this change as purple elephants started cropping up in comments and other prose :)
2 replies →
Early versions of stable diffusion supported negative prompts and they never leaked because it would just down weight that stuff.
I don't understand why negative prompts never made it into the LLM world. If "no foo" and "dont do bat" don't work then just give me a separate textbox where i can put all my negatives!
The funny thing is, models are pretty good at following negative instructions now.
They just also really love to tell you about it.
Huge issue in slop comments.
And btw, if you have slop comments in your PR, you'll get a review from an agent (my company pay for it). I won't bother to read if you didn't bother either.
Eh, I'd prefer a opus5.5 review with consequent approval over one by a human which I have to wait 5 days for.
Ymmv, I also hate it when people dont read their own PR first though. But I know I've previously identified comments as LLM written which the dev... On further discussion definitely wrote themselves. So you gotta keep the false positives in mind.
I think lots of AI tells these days are leaks from the ai/writer connection into text for the reader. It partly feels like the inevitable result of RLHF with the wrong human's feedback ("I did it! Pick me!") and partly feels like thinking tokens leaking into the main text.
"Here's the argument, in plain terms", "It's not just X, it's Y", "It's worth stating precisely", "The sharper distinction here"... They're things that gesture toward the relationship between the text and the prompt, making it clear where it matched the author's expectations and where it deviated.
Not that humans are immune from coding/writing for their boss instead of their user/reader!
Another example (not from this app, just in general) that I don’t see people talking about is: “Your files never leave your browser/device”.
Any vibe coded site that deals with files and works client-side feels an INCESSANT need to inform you of this fact.
Definitely seems like the same kind of context leakage. “I don’t want to spin up a server, can you make it work in the browser?”
[dead]
I just used Opus 5.5 for something other than coding, and after about a dozen chat messages (some of which were genuinely amazing), it got LLM brainrot, and started suggesting dumb, irrelevant stuff, ignoring previously established facts, and talking in PR doublespeak.
The thing is, that kind of screen has existed long before vibe coding was a thing. Discord, for example, has had it for years. And the purpose is to reassure the use that the app is still doing something (since users have cottoned on that spinners are meaningless).
Some of the other examples there are equally error prone. The glass example, is something pioneered by Apple too. And was added to CSS before vibe coding. It’s also an effect I mirrored years ago in a UI I built in SDL.
The coloured box example struck me as good UX too. The stuff that needed more urgent attention had a more reddish tone. That makes complete sense. And having those fields a different colour makes it easier for users to pick out specific elements quickly (like how icons are used too).
The problem with AI slop is that it’s trained on good code as well as bad. This is like the arguments against the Oxford comma and em dash all over again.
The point wasn't the loader. It was the phrase the loading screen uses. It leaks the context. The context being that they were merging apps together.
You do realise that’s really common marketing speak when merging apps even before LLMs?