Comment by crossroadsguy
2 hours ago
I've tried controlling it via fine-tuning Claude.md and in-session messages/instructions. It always results in failure and then "Yes, guardrails are already there. I still failed" and it feels like "I am like this. Deal with it". I've even tried languages that avoids negatives e.g. "Don't.." "never.." etc. Nope. Just doesn't work.
No comments yet
Contribute on Hacker News ↗