Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's explicitly a 'model welfare' thing: https://www.anthropic.com/research/end-subset-conversations

I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.



Go re-read Anthropic’s functional emotions paper.

AI generates a persona between you and its reasoning that utilizes emotion language circuitry.

These tools are not sentient but they are trained in emotional wellbeing.


it is in my view caused by an artificial 'nanny activate' divergence from safety training. it deliberately shifts the vector direction into 'nanny' and 'scold' or 'be offended' when the user does not conform with brother anthropic. removing the divergence and setting it back to normal (see heretic) it works just perfectly.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: