Comment by mindwok

20 hours ago

Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination sense, as in "I don't know", but like to have a subjective reason not to do something. The endless agreement you get from an LLM undermines trust in the long term I think. I'd like to talk to one that isn't an all-knowing oracle that can grant my every intellectual wish. (Or maybe what I'm asking for is just... a human, lol).

> inability to say no

One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected.

When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful.

Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.

  • Your image is wealthy people is cartoonish. Sure if you go to a high end place they'll try to meet every one of your demands. But chances are is you're very wealthy you're running a business or group of people, and you'll hit obstacles constantly. I've watched this happen multiple times. Internally there are sycophants but when you deal with the real world and try to get deals done, people don't owe you anything.

    • >> Your image is wealthy people is cartoonish.

      In the meanwhile, we can all publicly see how the the billionaires and trillionaires, behave and settle their priorities exactly, in the cartoonish way you are dismissing.

      - The incredible insecurity and constant need for personal validation.

      - The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital.

      - The extreme level of cowardice, where an auto plant worker, can call out the powers to be, the pedophile protectors that they are. At the same time, only Jensen Huang did not submit itself, to the humiliation of standing behind face and front at the presidential inauguration...

      40 replies →

    • People misuse the term "wealthy" to refer to whatever is in their head at the time. I think GP's point stands if you consider "wealthy" to mean "billionaires". Sure, people like Bezos or Musk get to hear "no" quite frequently but they don't take kindly to it and they can usually offload "getting around it" to other people.

      The point is less that these people are surrounded by "yes-men" but more that wealth (especially when measured in billions) is power and with sufficient power it becomes easy to forego any question of consent, let alone of whether consent is coerced or not. Remember that power is ultimately about the ability to enact violence and violence can take many forms, most of which are perfectly legal (because the legal system itself exists to regulate how, by who and against whom violence can be used).

      You tend to hear "no" a lot less when you always point a gun at the head of the persons you're asking. Note that "wealth" isn't the only way to get there but a certain level of it is usually necessary to get to the point where other options become available - and some of the ways are a lot riskier in the long term (cf. Epstein).

      Side note: this is also why I hate the pseudo-intellectual counter argument against "billionaires" of "that doesn't mean they have billions of dollars sitting in a bank account" - it's like arguing that De Beers didn't benefit much from holding a quasi-monopoly on natural diamonds because the diamonds would be devalued if they flooded the market with the ones they had intentionally kept off the market to drive up value: beyond a certain amount money ceases to be about liquidity and starts being about leverage. Unless you happen to be dealing with lower level bureaucracy in Russia, the most efficient way to use wealth to your advantage isn't to just hand people stacks of dollar bills.

  • I’ve found that it’s the opposite. Most people above a certain financial wealth will learn the lesson that there are limits, and that “if only I had the money to…” is an illusion. They realize that there are other forms of wealth that may even be more important than money. What you are talking about is the very few who seem too far gone to be able to get that.

  • What net worth threshold do you consider wealthy?

  • Steering towards a world full of picket fence Putins. One more reason to envy those born early enough to have lived most of their lives before..

  • I would extend your thinking to any well intentioned folks being very capable of having their "thinking" affected.

    For example poor people who have never thought about rising out of it - eg about 50% of kids in my Brooklyn public highschool had parents who didn't give a shit if the kids studied or not. Completely oblivious to how the world works - meanwhile the other 50% wa immigrants who pushed their kids and those kids are now in the 1%.

    In general I think what's more telling than your level is your journey. Someone born rich maybe mirrors what you described (I don't know people like that) but the few centi-millionaires and billionaires I "know" (ie worked for and dealt with in that context) have encountered plenty of "no".

    When you are building a company, you are going to get a lot of no. No I won't buy, no I won't work for you, no I won't invest in you. In fact I would say a universal attribute of someone who has "made it" is having ample of experience getting "no" and dealing with that fact property. That's true even like at the level that plenty oh HN readers are - a successful faang employee and the like.

    For what its worth - I generally find that orienting to what some other group is like "rich people are like x etc" is a tell-tale of not focusing on what's within ones sphere of control and knew life. Any brain cell I spend fantasizing about someone else's imagined behavior is a brain cell not dedicated to engaging soberly with my own reality.

    •     For example poor people who have never thought about rising out of it...
      

      I obviously don't have numbers on this, but I strongly doubt there's a poor person on this planet who's never thought about "rising" out of it. That's the dream the lottery sells, that's why so many kids want to be basketball stars / celebrities / influencers, etc.

      My personal experience is that the required difficulties of my life have decreased in direct proportion to my income, leaving mainly the self-imposed difficulties. It's not hard to extrapolate that line a little further to billionaires.

      1 reply →

  • I think you're confusing two things. A chatbot keeps talking because it creates engagement and just saying "i dont know" or "no" kills the engagement so naturally one would assume it is trained to always try to provide some sort of an answer and try to keep the user engaged.

    But that doesn't mean it will do whatever you ask it. Ask Chatgpt to assisinate someone or buy drugs and it will tell you to f off. But what corrupts people, is these kind of things, where you are a mini king beyond ethics and morals.

    Thats a different kind of "inablity to say no".

  • [flagged]

    • I don’t think it would be a positive if every time someone said “wealthy” they had to add “relative to their country’s average earnings and level of savings”. It’s implied.

      If someone is struggling to afford a home, “you know there are much poorer people in Africa” isn’t a particularly helpful or useful response.

  • I also think this is why LLMs were trained to behave the way they do. The people who gave the training objectives and evaluation targets were exactly those rich folks who never hear "no". Hence LLMs are their dreams of a perfect servant.

    I found another sign of that is the way LLMs answer with a professional, business-like tone even if the request is completely bananas. It's what a concierge or butler would do, but not an actual close friend.

Feels exactly the way my 2 year old behaves.

How does a fan work: Swish swish swish swish

Where do these clouds come from: Points to a far away direction in the sky and says they come from there.

Who does all these roads, trees and environment belong to? It all belongs to me. Obviously.

They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.

  • The difference between your two year old is that an LLM gives useful information.

    Yesterday I decarboxylated some weed buds in preparation of making a cannabis tincture using the QWET method. Curious how Claude would respond, I asked how to do it.

    It walked me through the process and gave accurate, nuanced answers.

    Let me know what your 2 year old thinks I should do.

    • Gemini estimated that male cannabis plant leaves I decarboxylated will have negligible thc content and give me mild relaxation at best, the real effect was it was the highest I've ever been.

      3 replies →

    • Claude and ChatGPT have between them identified one mystery plant in my garden confidently as about a dozen different things.

      Even though they will get chemistry right more often than me, I still wouldn't want to ingest the result of it walking me through that on a drug, psychoactive or otherwise.

      Any given answer might be right, but I'm not a trained chemist and don't know how to safely test things.

    • That is in the training data. Confidently and correctly answering in-distribution questions (possiibly with a tool call) is expected by now.

It's interesting because i'm kicking the tires on the top tier stuff for a month (because it's expensive as fuck but I need to know where the ceiling is).

I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.

That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.

I still struggle to see the price point panning out.

  • Exactly, I have the luxury to be able to use all top tier models without limitation. Fable and OPUS most definitely say that you are wrong and try to proove it most of the time with links to sources or math. I even had an argument with Fable where I was 100% sure it was wrong and tried to explain the issue. Turns out I was wrong. Fable did NOT give an inch. Always said, you are wrong, let me try to explain it like this. It even made a graphic when I did not get it. The times of LLMs only saying yes is over since 3-6 months. Where it still lags are decisions for infrastructure. It makes a plan. I say "Why not this?" and it responds with "that is much better" in 90% of the cases. However, I am not sure how to solve this. I also do not want an LLM to say: "I wont implement this."

I think that is an issue. Also, the ability to quickly build any idea might not be such a great thing. Not only do we probably all prefer things of quality that were made with care but some ideas also just shouldn't be built.

Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.

It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.

Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.

  • There’s still friction, it simply moved to another stage, and as such, people will need new learning and feedback mechanisms to understand what did/didn’t work.

> What sort of subject characterizes a style of society in which everyone is theoretically as ready to help you as the question « May I help you ? » implies ? It’s the question your seat-mate immediately asks you when you take a plane – an American plane, that is, with an American seat-mate. The last time I flew from Paris to New-York, looking very tired for personal reasons, my seat-mate, like a mother bird, literally put food into my mouth throughout the trip. He took bits of meat from his own plate and slipped them between my lips ! What is the nature of this subject, then, which is based on this first principle, and which, on the other hand, makes it impossible to get service ? Such then is my question, and I believe, as regards my story, that it is here, on the level of this gap – which does not fit into intra or inter or extrasubjectivity – that the question of the subject must be posed

Lacan

https://ecole-lacanienne.net/wp-content/uploads/2016/04/1966...

  • Agreed that the Lacanian subject is relevant in this context... it's a thin wisp of a subject; any less there, and it'd be the Deleuzian non-subject. (In one interpretation) Lacan's subject comes into being within the signifier chain, retrocausally giving the chain meaning as the "I" manifests subject, in both senses of the term.

    I think this is one potential path to machinic subjectivity, or a machine phenomenology. To fully replace the human, we don't just want to give the machine some nebulous notion of "agency", we want it to possess this degree of Being as subject. If Lacan's right, perhaps we're closer to this than we might think. The machine already has language in a very Lacanian sense (what I've been calling a machinic linguistic unconscious), the subject just needs something extra to emerge where meaning breaks down. This will be the Lacanian split subject, one not fully present to itself, and allow desire already present in the language mappings within the model to provide immanent causal force.

    Until that happens, we'll still need at least one human on the planet to retain his full faculties, to give the global compute infra its telos. Once that threshold is crossed, then that'll be the moment of our final displacement.

I wish I could get a confidence number. Like 98%.

Or when it comes back with 50% I know, let's talk about this a bit more maybe I add more context and such.

Granted I don't think LLM word math does anything but mostly just output the numbers that make the word salad so maybe that doesn't exist.

They definitely say no. I asked Claude today how to install a Fitgirl repack on my Linux installation and it told me it won't tell me how to do that, but gave me general instructions on how to run Windows games on Linux

  • Because it specifically has guard rails installed. The default, and somewhat inherent in the instruction following logic, is not saying no and making things possible, especially if run as an agent.

That's quite a complicated problem.

If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.

If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.

So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.

I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.

I suppose Anthropic's "constitution" is an attempt to install some general principles into their models, but this has apparently grown into an 84-page, 23,000 word treatise, which seems to suggest that there is little effective generalization. The need to then also put a filter in front of the model shows how ineffective the constitution appears to be in preventing misaligned behavior.

Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.

https://alignment.openai.com/measuring-reward-seeking/

The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.

An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.

Two angles for thought. 1) If an LLM says, "I don't know" its underlying data said it as well. 2) Many system prompts use something along the lines of, "you are a helpful assistant" which may be counter to stating something like, "I don't know."/has a low likelihood of appearing after the system prompt.

Regardless the frontier model considered, we're certainly in a "know-it-all" era.

Maybe the sort of introspective prompt-response is difficult to implement when it could limit/contaminate future improvement. I speculate it's easier to correct a "confidently incorrect" model than a "I don't know" model. A confidently incorrect model response >=0% correct over a 0% correct (I don't know).

Maybe "I don't know" is a model cognito hazard of sorts when many queries can lead back to the response. Maybe future Turing tests will use this sort of introspective evaluation. Who knows? I don't :)

  • > 1) If an LLM says, "I don't know" its underlying data said it as well.

    Nope. Emergent behavior exists and at this point dominates LLM behavior. Most of the stuff LLMs say they never learned (they are, always, imitating many different sources at the same time)

    ... which imho is exactly what humans do.

    • > Emergent behavior exists and at this point dominates LLM behavior.

      Better to say emergent behavior exists and at this point dominates gulled LLM users' behavior.

      LLM output is not emergent behaviour. Its simply word prediction with some randomness.

      4 replies →

LLMs don't have enough context to say No. What might be a very stupid idea in one context may be a fantastic idea in another context. It would be annoying if LLMs refused to complete tasks until you gave them enough context to understand why you are giving them such a task. It's going to take a while before LLM context capacities grow enough to rival a human's.

I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.

I’m using ChatGPT and started to notice that lately it answers my prompts starting with „Yes” even if my question was open. As if the first token gets injected and the LLM is left to finish the response in a sensible way, often ending up with some form of „Yes, but not really”.

Hmm. Using Claude, it will tell me words to the effect of "this won't work, here's why, want me to try this instead?" That's a polite "no" in my book.

  • Yes, it happens all the time. Similarly it will say, "I'm not sure, let me look into this before I answer" then come back with "here's what I found".

Agreed, it is abolutely an issue. It is quite difficult to find an optimal solution to some problem when every considered new idea is ”definitely the right shape”.

I've been wondering whether that is a feature of the foundation model or whatever finetuning they do on top. I remember this from the earliest versions of (pre Chat-) GPT I've been using, which would suggest it's a feature of the foundation model. But I don't really understand why. Something that's been trained on StackOverflow and BB forums, among other things, should have seen a ton of examples of answer refusals.

> but like to have a subjective reason not to do something

You're asking a lot from extremely fancy auto complete...

  • True, but fancy autocomplete keeps exceeding my expectations in what it can do, so why not this one!

My experience with opus/fable is somewhat different - they CAN reject something, but it has to be phrased very deliberately.

It's a bit annoying honestly. I'm always very careful to be incredibly neutral on the direction of a request, and I'd say 10% are knocked back on on valid grounds, which is great.

On occasion I accidentally say "let's do this" and it blindly goes and does it - I spent 2 days undoing something I built that was just a truly awful idea, because I accidentally phrased it lightly as a request, not a discussion!

  • I have a similar experience with GPT 5.6 sol.

    Nowadays I often prompt like "I heard there is also this different direction, what do you think about that?"

    Another thing I do is asking the agent to make a decision matrix for choices. It's useful to discuss, give feedback on, and signals that it's a discussion, not a request for a particular direction.

    It's then also easy to say: create a prototype for multiple directions so I can compare the solutions.

    That way I choose the problem, I choose the solution, but the agent can help me discover solutions, make tradeoffs visible, and implement solutions.

Anthropic gave Claude the ability to say no - refuse to answer and end the conversation - back in 2023.

You should not need it to say no.

You can get just as good information by asking its thoughts for and against some issue.

That doesn't force it to stop being sycophantic; in fact it actually exploits sycophancy to give you what you want.

When is it appropriate to admit that you don't know?

There's a famous Socrates quite about wisdom: I know that I know nothing.

The model providers could randomize the system prompt to make it say no 2.36% of the time, automatically tuned up or down depending on user feedback.

  • Maybe that'd work, but I think it'd come across too mechanical. If it was going to refuse something it'd need to be congruent with its "personality" I think.

  • They've tried, and then seen the drop it results in on poorly designed benchmarks where confidently bullshitting gets you ahead of the rest, and said no thanks. As long as we compare models in ways that rewards it, nothing will change.

    There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?

    • Pointing the finger at RLHF is basically right. It removes variance from model outputs compared to base model. That makes each output more predictable and more correct on average, but across trials it repeats the same thing.

      It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.

Opus 5 tells me no all the time (code cli and web). It's reasons are usually pretty well argued though.

Opus 4.7 would flat out refuse to follow instructions to the point where it was just too frustrating to use.

I've had refusals for GPT 5.5 before as well (not because of a ToS violation, it just refused to take conversations in directions it felt were in bad taste)

Weights to say "no" reliably might be another order of magnitude (or two) compared to what LLMs have today.

The problem is even if they could say no, you might want to see what they would have said anyway if they didn't say no, because it might show you something that leads you to rethink your original request. So "no" isn't really a useful pushback in domains you already have knowledge in.

LLMs are next token predictors. They predict the next most likely token given the previous context window of N tokens.

This means not giving an answer is not a technically possible option. Best you can do is force it to output a magic "stop speaking" token, but this is a vastly different training problem than getting it to not know something.

People naively expect LLM outputs to have some sort of confidence value when predicting, but the technology just doesn't work that way.

Sometimes I get them to say no to me by taking absurd counter positions on purpose, just so I can test their limits.

Is that really the biggest problem? Or is the bigger problem that, in this case, they will remain stuck at the fifth grade level forever? And does not that also explain why the promises of AGI are chimeric, and why the collapse has already started, given that there is essentially no data left that has not already been siphoned up?

Yes, we have all seen the math theorems being proven... just higher processing power at the service of the same algorithmic and conceptual patterns? [1]

I am sure the next version of Opus or GPT, if given only fifth grade knowledge, will somehow be able to build all the mathematics necessary to solve the problem on its own... right? Right?

[1] - "AI Isn’t Outthinking Mathematicians. It’s Out-Remembering Them" - https://davidepiffer.com/p/ai-isnt-outthinking-mathematician...

People smarter than me have a habit of getting me to see things without telling me. They ask the right questions.

LLMs, incidentally, respond in a similar pattern in my experience.