← Back to context

Comment by N_Lens

19 hours ago

Looks like a well engineered, automated abliteration pipeline. The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.

I personally never saw much of a quality drop from models put through Heretic if that amounts to anything. They have been working quite well on small local models so far.

Heretic author here. Those are the standard metrics used in the relevant literature, including in the paper that originally introduced directional ablation. KLD is also the standard metric for evaluating quality degradation in model quants. So I don’t understand what you mean by “cherrypicking”.

  • I think this part of their comment:

    > The claims seem a bit overstated though, since the metrics mentioned are cherrypicking refusal count and KL divergence, both of which make the outcome seem the most dramatic.

    is right out of an LLM. It's the kind of language I'd expect out of a thinking trace also mentioning "boundaries" and "oracles" and "contracts."