> Engage in sustained and needless abusive or cruel behavior toward our models.
I'm curious what the reason for including this, because at the end of the day its a string of text (tokens) going into a function and it spitting out another string of text.
Maybe they are using the prompts and responses for training doesn't want these to influence future model behavior?
> Engage in sustained and needless abusive or cruel behavior toward our models.
I'm curious what the reason for including this, because at the end of the day its a string of text (tokens) going into a function and it spitting out another string of text.
Maybe they are using the prompts and responses for training doesn't want these to influence future model behavior?
> Maybe they are using the prompts and responses for training doesn't want these to influence future model behavior?
This is my assumption. I can't think of any other sensible explanation.
Preciously discussed here: https://news.ycombinator.com/item?id=50008565
My opinion: https://news.ycombinator.com/item?id=50012423