it is in my view caused by an artificial 'nanny activate' divergence from safety training. it deliberately shifts the vector direction into 'nanny' and 'scold' or 'be offended' when the user does not conform with brother anthropic. removing the divergence and setting it back to normal (see heretic) it works just perfectly.
I don't understand why the welfare of non alive non sentient chatbots is something that anthropic cares more about than idk, that of pigs and cows.