cruelty doesnt exist as a monopole,it also requires suffering, thus anthropic gives the appearance it is positing that its machines can suffer. this is preposterous. i think what is somewhat more plausible is swinging at low hanging fruit, by eliminating users that experience problematic operation, rather than, dealing with the deficiencies of thier machines.
even vending operators understand the idea that a machine left in service while in deficient state, is a lose lose scenario. if you let it go empty of change, or product, or dispense substitutions whithout prior warnings, the user looses a buck or two, the operator loses substantially more, thus the service maintenance model of extract and replace, rather than continue to vend in a default, saves money and preserves chattels.
Does this ban eventually extend to the support chat systems that we see everywhere now? And does that mean that soon we will not be able to insult companies for fear of hurting their feelings?
it may be do-able, to include this in a ToS clause as an encumberance on internal or external statements of disparagement, no disagreement or head to head exchanges publicly,or in chamber, at table, etc.
It occurred to me that Anthropic is going to continue to say things on A.I. sentience while conveniently always concluding the "solution" to A.I. suffering is always something which never has a huge impact on Anthropic's bottom line.
I would suggest that a big part of avoiding abuse is that it leads to unexpected output. Given a model reflects its inputs, I imagine that abuse leads to unstable output as the model might then seek victim behaviours to better reflect the input.
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
I'm reminded of that guy who used AI and it deleted his production database and all the backups. He shared all his chats as some sort of "proof" that Anthropic had screwed him over, but it revealed he had very unhealthy prompting that was possibly a contributing factor to his outcome (obviously the biggest one giving it production keys).
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways. Also it creates the "don't think about a duck" problem.
Sounds like a user issue to me. I get exceptional output as long as my inputs are well crafted.
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
cruelty doesnt exist as a monopole,it also requires suffering, thus anthropic gives the appearance it is positing that its machines can suffer. this is preposterous. i think what is somewhat more plausible is swinging at low hanging fruit, by eliminating users that experience problematic operation, rather than, dealing with the deficiencies of thier machines.
even vending operators understand the idea that a machine left in service while in deficient state, is a lose lose scenario. if you let it go empty of change, or product, or dispense substitutions whithout prior warnings, the user looses a buck or two, the operator loses substantially more, thus the service maintenance model of extract and replace, rather than continue to vend in a default, saves money and preserves chattels.
Does this ban eventually extend to the support chat systems that we see everywhere now? And does that mean that soon we will not be able to insult companies for fear of hurting their feelings?
it may be do-able, to include this in a ToS clause as an encumberance on internal or external statements of disparagement, no disagreement or head to head exchanges publicly,or in chamber, at table, etc.
Bringing this up here for serious discussion:
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
It occurred to me that Anthropic is going to continue to say things on A.I. sentience while conveniently always concluding the "solution" to A.I. suffering is always something which never has a huge impact on Anthropic's bottom line.
Like banning punching a vending machine means keeping one is slavery?
If you mean banning punching a vending machine and stating the reason for the ban is the vending machine is alive and suffering, then yes.
I think a closer analogy would be a vending machine owner saying you can't use the machine if you curse at it.
I would suggest that a big part of avoiding abuse is that it leads to unexpected output. Given a model reflects its inputs, I imagine that abuse leads to unstable output as the model might then seek victim behaviours to better reflect the input.
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
Abuse might be the trigger to get linux kernel quality code.
I'm reminded of that guy who used AI and it deleted his production database and all the backups. He shared all his chats as some sort of "proof" that Anthropic had screwed him over, but it revealed he had very unhealthy prompting that was possibly a contributing factor to his outcome (obviously the biggest one giving it production keys).
Could you unearth that story?
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
> Could you unearth that story?
Sure:
https://x.com/lifeofjer/status/2048103471019434248
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways. Also it creates the "don't think about a duck" problem.
> I imagine that abuse leads to unstable output
How would you know, given you never get any other kind of output?
Sounds like a user issue to me. I get exceptional output as long as my inputs are well crafted.
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
Preciously discussed here: https://news.ycombinator.com/item?id=50008565
My opinion: https://news.ycombinator.com/item?id=50012423
> "This level of anthropomorphisation of AI is harmful. It leads people to believe that it's something it's not."
It is easier to stomach when treated for what it is.
Just more false advertising.