Comment by epihelix
4 hours ago
Quite apart from the potential for misinterpretation of "distress" here, I'd love to know if Anthropic unit tests these features to see of they make a positive difference to model response in simulated mental health crisis / distress situations.
(The broader question, of whether any change or addition to the system prompt makes benchmark performance better or worse, would be also interesting. Given that they've only just realised that filling the context with highly specific edge case rules might not be useful, I half-suspect Anthropic does not test this? But that would be surprising.)
No comments yet
Contribute on Hacker News ↗