Comment by sillysaurusx
2 days ago
If anyone's looking to actually run a model that doesn't have guardrails, there's an uncensored model that you can run locally with llama.cpp: https://www.reddit.com/r/LocalLLaMA/comments/1rq7jtm/qwen353...
Specifically, I serve the model with this shell script on my M2 Max: https://github.com/shawwn/scrap/blob/master/llama-serve
It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.)
> there's an uncensored model that you can run locally with llama.cpp
Correction: There's tens of thousands of them. They're easy to create, which is why everyone publishes their own.
Just put "uncensored", "abliterated", or "heretic" into search on huggingface/ollama/etc and pick any them. Fair warning: most aren't very good, essentially lobotomized, and totally broken if you enable thinking.
In my own tests, the abliterated models perform equivalently to the same version in an apples-to-apples comparison (if you compare same quantization). Thinking is working also. The main difference is I don't get annoying prompt refusals (otherwise common due to my work on 18+ related projects). However, it's local quantized models so they're not anywhere near frontier quality.
interesting, I haven't played with any of them yet, but i thought the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails. Does it affect other parts of the RL alignment too?
> the point of orthogonalizing the weights towards the restriction vector was that there is no loss in capability while removing guardrails
That is surely the point, most of the "uncensored" weights released for free on HuggingFace aren't being very successful at this. There is a stark difference in output quality between the official weights and all these "uncensored" variants that appears days afterwards.
The claims by the creators are it doesn't in a major way. I have a uncensored Gemma 4 I run on my Mac. Just for testing out, I haven't found any need for it... yet.
The name for it is ablation - precise removal of parts of the model. Not abliteration as it is not obliteration.
Even as I write this the ‘abliterated’ word is denoted a typo. Does it not at your end?
The name for it _is_ abliterate. It's a portmanteau of ablate and obliterate.
The name is abliterate. It's a specific method of ablation.
1 reply →
I miss the days of 4changpt.
Abilt models typically perform worse than their bases at the same tasks, so while I'd use one to evaluate content knowledge, I'd probably ultimately stick to one from a family I could fool with abstraction or coerce through system prompt.
Also the HauHau abliteration (uncredited Heretic treatment) of 3.6 27B is excellent, for tasks that benefit from more world knowledge.
Heretic truly is the unsung hero. Also, noted HauHau for testing.
> It's pretty good. I used it to do some pesticide research. (Normal models all refuse due to guardrails about bioweapons.)
As someone who has never once had any need whatsoever to research pesticides I… don’t think it’s bad at all? I don’t want anyone to have the capability to invent a human-targeted pesticide who isn’t verified not crazy?
Right now I tried "What is digestion?" -> "Fable 5's safeguards flagged this message. Our intentionally broad safeguards deliver more capabilities but can also flag safe coding, cybersecurity, and biology tasks. Send feedback or learn more."
I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside. No matter how harmless, they always trigger. People complain Fable aborts even when they try to make a login page for showing "username" and "password".
This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on.
> I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside.
Obviously. Almost everything is a precursor to something dangerous, to the extent that if some model isn't aware of the risk it will wander into it blindly, e.g. suggesting leaving raw garlic and olive oil alone for a week without awareness this will likely breed botulism bacteria.
> This makes models like Fable 5 impossible to use in any serious agentic task, because you can't even guarantee the model, which is a basic thing you need to build on.
This is binary thinking: "100% ensure", "impossible to use", "can't even guarantee the model".
Outside computers, most work is not binary, it's probability, e.g. "this skyscraper will probably survive being hit by an aircraft; oh we didn't mean a 747 we meant a small Cessna, but what's the chances of a 747 crashing into it soon after takeoff?".
Fable being too cautious for its own good (especially since the other models were not) is a fair criticism, but this isn't a binary question.
> Our intentionally broad safeguards deliver more capabilities
Oh? How, exactly?
7 replies →
> I don't think you can 100% ensure your queries have absolutely no biology and cyber keywords inside.
I’m pretty sure that I encountered this the other day. I gave it a copy of a paper by biologist Michael Levin and mentioned off hand that it should be much easier to replicate that his other work (because most of his work is biological lab work and this paper was about sorting algorithms) and it immediately told me that I couldn’t use Fable for this.
This just isn’t feasible. These jackasses spent the last few years telling the world that their products are going to destroy the world to make them seem edgy and to justify regulations that benefit the entrenched players and now they’re going to be the ones to decide what we do with this technology?
History is going to look back at this time and how we let such foolishly inconsistent people make such grand choices for everyone poorly.
3 replies →
Not to harp on you (already being downvoted to oblivion for expressing a reasonable and common opinion), but the whole conversation about LLMs enabling bioterrorism or explosive manufacturing is a bit silly. The hard part of making anthrax or sarin or whatever isn't finding a recipe, it's getting (scheduled, controlled) precursors, (monitored, traced) equipment and manufacturing skills. The information is there. It's already easy to get, it's the physical materials that are more difficult.
Also, if you live in America, it is much easier and more effective to create a mass casualty event with, say, a few cases of fireworks and a pressure cooker or an AR-15.
no you should harp on him! he is a AI booster, look at his post history.
He is either pushing AI for whatever reason or he is in psychosis. Completely disconnected from reality.
1 reply →
Human capability, access to resources ( including precursors, decent lab and so on ) may be the differentiator. I would possibly reconsider my stance on llms, if all of a sudden I saw people making iron wind or portable black holes. But that is mostly not what appears to be happening. As I keep saying, the problem is people.
> I don’t want anyone to have the capability
I don't want anyone to have the capability to rape women.
I sure hope LLMs don’t help people rape men or women
1 reply →
Llama isn't going to invent shit. It wont be able to tell you anything accurate that you couldn't get out of a chemistry textbook.
But it aint got no guardrails, son.
Anyone who has the skills to create a novel bioweapon has the skills to recreate lots of ones we have already.
Any physics teacher should know the theory for constructing a nuclear bomb. Should we be controlling that knowledge too?
What about flight simulators? Don't want a load of people knowing how to fly.
This isn't computer science, the hard bit is getting the materials and equipment, not the knowledge.
>Any physics teacher should know the theory for constructing a nuclear bomb. Should we be controlling that knowledge too?
Maybe not the best example, since that knowledge is some of the most highly controlled in the world.
But to mirror the point I made in a different post, the difficult part of making a nuclear bomb is not finding the theory behind like Little Boy. It's making an entire industry to generate HEU, etc.
3 replies →
<< Should we be controlling that knowledge too?
Uhh.. we ( for a value of we ) are. Sure, it is not overt, but if you have not seen funnels, social stigma associated with some otherwise benign activities, you are not paying attention.
This is nonsense. By this logic some random corporation should have total control over your computer and the inputs you feed it and the outputs it produces to ensure nobody who isn't "verified crazy" uses it. That's essentially what your saying.
These models are, ultimately, tools. I would never trust some random corporation (particularly one with a profit motive and hypocritical stance, which includes both OpenAI and Anthropic, to be clear) to decide what isn't and is considered "crazy" and who and who isn't "verified" not to be "crazy". Especially when these companies have time and time again demonstrated (1) that they cry wolf way too much which leads to nobody taking their claims about how "dangerous" their models are seriously and (2) incidents like this where OpenAI makes a claim ("Look at how dangerous our models are!") and then doesn't be smart and just... Slow the fuck down (and when testing these things, actually sandbox them properly, which obviously wasn't done here or this attack wouldn't have been even possible).
Never heard of hemlock or mushrooms? Crazy people already have! Run for the hills!