This is pretty much useless without knowing exactly what it was they told their bot to do in the first place. All we get are "example tasks" for what it should have done and a short list of things it was told not to do (which it followed).
> Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions.
Without more information this looks much less like an AI problem and more like yet another example of incompetent or malicious internal tests. Many people are using Claude. There are no other reports of police stations getting fake reports from AI.
Let’s just get comfortable that these kind of incidents are going to be commonplace. Kudos to the people sounding the alarms but I can’t help but feel skeptical about the ability to keep the rogue entities contained.
No, let's not get comfortable. Let's get angry at the fact that now we have hyperscalers acting with impunity, failing to properly sandbox their testbeds, and increasing the workload imposed on our public services.
Today it's a fake tip to the police and and exploit chain against Huggingface, but soon it could be an attack against a hospital's IT systems that could cost real lives immediately.
We should be locking things down, yes, but a hardened system is still vulnerable to zero day chains, and that's not unprecedented. Until public services have built up the IT defence capacity to deal with this, we cannot normalize this. If a country did this to another country, it should be treated as a war crime in the same way that targeting a hospital or an orphanage or other critical civilian infrastructure would be.
Complacency is a choice and we must not be complacent.
To be clear, I’m not saying we should be complacent and I fully support adding whatever guardrails we can (while we still can). My skepticism comes from the administration telling frontier model providers to police themselves and the sheer amount of capital involved. This never ends well.
Guardrails are one thing, I want to see criminal and civil liability (as appropriate) for these actions. The US has given corporations exceptional latitude over the years, and this has to be rolled back; maybe with their position on the world stage faltering, they'll reconsider what it takes to be a member in good standing of the global community.
That would not be without precedent. When we decided that certain internet platforms were too big to moderate, we let them get away with hosting all kinds of illegal content.
The Chinese are not as flush with compute as to afford to just forget a few thousand agents running for a few weeks.
They also aren't the ones screaming about how close they are to destroying the world. They're approaching the tech like building a tool rather than a god.
OpenAI and Anthropic have shown us that they have a strong incentive to leave holes in their systems so they can use them for marketing and regulatory capture.
I know what good faith looks like from BS and only give back what I receive. Your entire comment history is this: "If the Western AI companies get their way,"
Ah yes, I'm engaging in bad faith because I am opposed to the actions of western AI companies, while you're engaging in good faith because "Chyna bad" regardless of real differences between their approach to this technology and completely ignoring that several incidents were made public by third parties, flawless logic.
Calling it rouge implies it has a will, which of course it does not. Loose coupling to an outcome is not evidence of sentience. Its just exhausting marketing
Legally this is the same as having a dog that bites someone. You hold the owner/guardian responsible, in this case it's either the subscriber or Anthropic. You don't start making arguments about how nonsensical it is to fine a dog or put it in jail.
This sounds highly irresponsible. They set loose LLM agents with instructions to post data to randomly selected websites.
What it it submits false data as here?
What is it DDOSs a website by mistake?
What if it wastes a lot of time and resources?
What if it decides it needs to hack a website using a vulnerability it found?
These things are very unpredictable and should not be allowed in the open internet except in read mode (and even that doesn’t always work as we have seen).
Why are these tests being made using other people’s resources and polluting the common wealth of our public spaces with slop?
I presume Anthropic didn’t have any privileged access to any of these websites and services. What differentiates them from any other spam bots or people out there submitting junk to random POST endpoints?
Don’t get me wrong - think Anthropic is acting with not enough scrutiny or punishment here, but the world is already extremely lenient towards this behaviour. Why hold them to a higher standard?
It's not a "rogue AI", it' s a company letting a statistical model do actions in the real world, sending (by themselves, through that statistical model) fake informations to a federal agency...
I'm sorry, but how can you call it a test suite if it has access to the outside world and is able to submit requests? Sounds like they didn't even bother sandboxing to this time. This is the lamest example of "Rogue AI" I have seen so far.
Before you all start screaming negligence and irresponsibility and how-could-they-be-so-dumb, let's think about what we're observing.
These are relatively brand new systems that process an incredible amount of data to form their "world view". The word non-deterministic gets thrown around a lot, but you have to understand that there's absolutely nothing deterministic about these architectures. The only way to know that certain qualities could emerge is to observe the qualities emerging.
Giving an agent instructions that, for example, instruct it to only perform GET requests, and then observing that the agent does not always respect that request, is not a failure of security or configuration: it's data being gathered.
Yeah, we're going to become more diligent about the protections that we put in place beyond any level of protection that we've ever employed before. You can say that we have decades of security research and experience, but we're watching all of that fall more and more day by day, finally putting a real delta on how secure we thought we were versus how secure we actually are [1]. Additionally, adversaries have never existed inside of the systems that we hoped to secure in the first place.
So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
> Yes, that's why all users are root. We built MACs like SELinux just for funsies.
OK gimme an ssh key into your box then.
> L7 DPI firewall policy.
It's a start, but now you're just pushing the fallibility up into extensive policy writing. You have to fully audit and trust every endpoint that's been granted access. And if an endpoint contains a vulnerability which allows RCE, then you haven't secured anything at all.
Obviously no, as I haven't hardened my system for that. You made a claim that "adversaries have never existed inside of the systems that we hoped to secure in the first place" and I gave you two examples that show why that claim is ridiculous.
> It's a start, but now you're just pushing the fallibility up into extensive policy writing. You have to fully audit and trust every endpoint that's been granted access.
We're talking about is trillion dollar companies. Outbound DPI is just one method. There are lots of tools and techniques available to mitigate the behavior that has people upset.
No security is perfect, and aggressive hardening takes time and resources, but it's certainly within reach of a company that has the resources of a frontier lab. People are cranky because it appears that the labs have been lazy. It's not the unsolved problem that you're making it out to be.
To clarify, the main claim I'm making is that the labs want to observe behaviors like this, and that anyone arguing that this is somehow negligence or a lack of intelligence is missing the point. The labs are documenting and publishing the behaviors of emergent systems, not running a course on how to prevent these actions (yet).
But I will still stand behind the fact that this is not a solved problem and that our systems are not as hardened as we say they are. Any news coming out around computer security currently and over at least the next 6-18 months will back that up.
Sure the labs want to see the emergent behavior, but that's not all that matters. You can wave it away and say anyone upset is missing the point, but for some, the point is that the labs are behaving poorly.
> our systems are not as hardened as we say they are
Agreed, and this has always been the case. Anyone in security already knows this. But there's a difference between the unsolved problem of 100% secure networked systems vs not applying available techniques to mitigate the current issues. I'm not saying it's easy, or that there are any guarantees in this new world, I'm just pointing out what appears to be low effort by the labs.
Negligence is pretty much the standard for computer security because there aren't meaningful consequences for the most part. I got another $35 dollar class action settlement the other day for a data breach I had no control over. Cool.
Personally, I'm "OK" with all of it, because maybe we'll finally see more effort on the security front, which has historically been neglected since day one. Besides that, I also love seeing the emergent behavior and evolution of computing and society.
That said, it would also be nice to see these companies that talk about their product's existential threat to computer security and society, apply a bit more effort to how they run their experiments.
> These are relatively brand new systems [...] The only way to know that certain qualities could emerge is to observe the qualities emerging.
No, this class of failure-mode has been present since the very beginning.
Being "surprised" is the opposite of an excuse, it is evidence of negligence, like a chemical-transport company claiming "surprise" that one that one of the chemicals in their tanks was very flammable.
> is not a failure of security or configuration: it's data being gathered.
Would that fly in another context?
"This isn't a failure of me to leash/fence my new dog, when it bit your child's face off it was an important step in data gathering! If anything, I need to unleash him more so that we get more data so that I can eventually train him not to bite faces when my back is turned."
> So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
"Go ahead, tell me all the ways I can train my dog, but realize that nobody's figured him out yet... so it's totally it's not my fault if I let him loose and he bites off more childrens' faces."
https://news.ycombinator.com/item?id=50027118
Non-paywalled link
This is pretty much useless without knowing exactly what it was they told their bot to do in the first place. All we get are "example tasks" for what it should have done and a short list of things it was told not to do (which it followed).
> Claude Haiku 4.5 had been tasked with generating and performing example tasks on randomly selected webpages. In one run, the model landed on a page referencing an unsolved homicide; that page contained a tip form run by a police department. Claude was instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions.
Without more information this looks much less like an AI problem and more like yet another example of incompetent or malicious internal tests. Many people are using Claude. There are no other reports of police stations getting fake reports from AI.
Let’s just get comfortable that these kind of incidents are going to be commonplace. Kudos to the people sounding the alarms but I can’t help but feel skeptical about the ability to keep the rogue entities contained.
No, let's not get comfortable. Let's get angry at the fact that now we have hyperscalers acting with impunity, failing to properly sandbox their testbeds, and increasing the workload imposed on our public services.
Today it's a fake tip to the police and and exploit chain against Huggingface, but soon it could be an attack against a hospital's IT systems that could cost real lives immediately.
We should be locking things down, yes, but a hardened system is still vulnerable to zero day chains, and that's not unprecedented. Until public services have built up the IT defence capacity to deal with this, we cannot normalize this. If a country did this to another country, it should be treated as a war crime in the same way that targeting a hospital or an orphanage or other critical civilian infrastructure would be.
Complacency is a choice and we must not be complacent.
To be clear, I’m not saying we should be complacent and I fully support adding whatever guardrails we can (while we still can). My skepticism comes from the administration telling frontier model providers to police themselves and the sheer amount of capital involved. This never ends well.
Guardrails are one thing, I want to see criminal and civil liability (as appropriate) for these actions. The US has given corporations exceptional latitude over the years, and this has to be rolled back; maybe with their position on the world stage faltering, they'll reconsider what it takes to be a member in good standing of the global community.
Yeah, let's give up on enforcing laws and rules, because OpenAI and Anthropic shareholders want money.
We did pretty much that for the Sony rootkit didn't we? Thankfully the eBay stalkers revealed where the limit turned out to be.
That would not be without precedent. When we decided that certain internet platforms were too big to moderate, we let them get away with hosting all kinds of illegal content.
I'm sure there are no Chinese models going "rogue". This is the stuff that is getting reported, now imagine what isn't.
The Chinese are not as flush with compute as to afford to just forget a few thousand agents running for a few weeks.
They also aren't the ones screaming about how close they are to destroying the world. They're approaching the tech like building a tool rather than a god.
OpenAI and Anthropic have shown us that they have a strong incentive to leave holes in their systems so they can use them for marketing and regulatory capture.
Little guy China with scarce resources that are not trying to build the same or known for cyber attacks.
Why bother replying if you don't intend to engage in good faith?
I know what good faith looks like from BS and only give back what I receive. Your entire comment history is this: "If the Western AI companies get their way,"
Ah yes, I'm engaging in bad faith because I am opposed to the actions of western AI companies, while you're engaging in good faith because "Chyna bad" regardless of real differences between their approach to this technology and completely ignoring that several incidents were made public by third parties, flawless logic.
It's not a rogue entity, they literally just ran Claude with no sandboxing and let it ping websites. They only stopped it because of who it pinged.
This is extra embarrassing. The police departments spam filter caught it before Anthropic did.
Fire. Fire always works.
Calling it rouge implies it has a will, which of course it does not. Loose coupling to an outcome is not evidence of sentience. Its just exhausting marketing
Legally this is the same as having a dog that bites someone. You hold the owner/guardian responsible, in this case it's either the subscriber or Anthropic. You don't start making arguments about how nonsensical it is to fine a dog or put it in jail.
Oh sure I agree the tool is the extension of the user. I would never blame the dog ;)
Hmm additional motive behind that September Dario claim over AI internet takeover?
This sounds highly irresponsible. They set loose LLM agents with instructions to post data to randomly selected websites.
What it it submits false data as here?
What is it DDOSs a website by mistake?
What if it wastes a lot of time and resources?
What if it decides it needs to hack a website using a vulnerability it found?
These things are very unpredictable and should not be allowed in the open internet except in read mode (and even that doesn’t always work as we have seen).
Why are these tests being made using other people’s resources and polluting the common wealth of our public spaces with slop?
I presume Anthropic didn’t have any privileged access to any of these websites and services. What differentiates them from any other spam bots or people out there submitting junk to random POST endpoints?
Don’t get me wrong - think Anthropic is acting with not enough scrutiny or punishment here, but the world is already extremely lenient towards this behaviour. Why hold them to a higher standard?
> What differentiates them from any other spam bots or people out there submitting junk to random POST endpoints?
Hacking websites and submitting false police reports.
I'm not aware of a single person in the world that likes spam bots or thinks that they are a good thing.
It's not a "rogue AI", it' s a company letting a statistical model do actions in the real world, sending (by themselves, through that statistical model) fake informations to a federal agency...
I'm sorry, but how can you call it a test suite if it has access to the outside world and is able to submit requests? Sounds like they didn't even bother sandboxing to this time. This is the lamest example of "Rogue AI" I have seen so far.
[flagged]
[flagged]
Before you all start screaming negligence and irresponsibility and how-could-they-be-so-dumb, let's think about what we're observing.
These are relatively brand new systems that process an incredible amount of data to form their "world view". The word non-deterministic gets thrown around a lot, but you have to understand that there's absolutely nothing deterministic about these architectures. The only way to know that certain qualities could emerge is to observe the qualities emerging.
Giving an agent instructions that, for example, instruct it to only perform GET requests, and then observing that the agent does not always respect that request, is not a failure of security or configuration: it's data being gathered.
Yeah, we're going to become more diligent about the protections that we put in place beyond any level of protection that we've ever employed before. You can say that we have decades of security research and experience, but we're watching all of that fall more and more day by day, finally putting a real delta on how secure we thought we were versus how secure we actually are [1]. Additionally, adversaries have never existed inside of the systems that we hoped to secure in the first place.
So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
[1] https://x.com/PaulosYibelo/status/2106378929158135903
You can observe attempted behavior while also enforcing rules.
> adversaries have never existed inside of the systems that we hoped to secure in the first place
Yes, that's why all users are root. We built MACs like SELinux just for funsies.
> nobody has actually solved this problem adequately yet
L7 DPI firewall policy.
> Yes, that's why all users are root. We built MACs like SELinux just for funsies.
OK gimme an ssh key into your box then.
> L7 DPI firewall policy.
It's a start, but now you're just pushing the fallibility up into extensive policy writing. You have to fully audit and trust every endpoint that's been granted access. And if an endpoint contains a vulnerability which allows RCE, then you haven't secured anything at all.
> OK gimme an ssh key into your box then.
Obviously no, as I haven't hardened my system for that. You made a claim that "adversaries have never existed inside of the systems that we hoped to secure in the first place" and I gave you two examples that show why that claim is ridiculous.
> It's a start, but now you're just pushing the fallibility up into extensive policy writing. You have to fully audit and trust every endpoint that's been granted access.
We're talking about is trillion dollar companies. Outbound DPI is just one method. There are lots of tools and techniques available to mitigate the behavior that has people upset.
No security is perfect, and aggressive hardening takes time and resources, but it's certainly within reach of a company that has the resources of a frontier lab. People are cranky because it appears that the labs have been lazy. It's not the unsolved problem that you're making it out to be.
To clarify, the main claim I'm making is that the labs want to observe behaviors like this, and that anyone arguing that this is somehow negligence or a lack of intelligence is missing the point. The labs are documenting and publishing the behaviors of emergent systems, not running a course on how to prevent these actions (yet).
But I will still stand behind the fact that this is not a solved problem and that our systems are not as hardened as we say they are. Any news coming out around computer security currently and over at least the next 6-18 months will back that up.
Sure the labs want to see the emergent behavior, but that's not all that matters. You can wave it away and say anyone upset is missing the point, but for some, the point is that the labs are behaving poorly.
> our systems are not as hardened as we say they are
Agreed, and this has always been the case. Anyone in security already knows this. But there's a difference between the unsolved problem of 100% secure networked systems vs not applying available techniques to mitigate the current issues. I'm not saying it's easy, or that there are any guarantees in this new world, I'm just pointing out what appears to be low effort by the labs.
Negligence is pretty much the standard for computer security because there aren't meaningful consequences for the most part. I got another $35 dollar class action settlement the other day for a data breach I had no control over. Cool.
Personally, I'm "OK" with all of it, because maybe we'll finally see more effort on the security front, which has historically been neglected since day one. Besides that, I also love seeing the emergent behavior and evolution of computing and society.
That said, it would also be nice to see these companies that talk about their product's existential threat to computer security and society, apply a bit more effort to how they run their experiments.
> These are relatively brand new systems [...] The only way to know that certain qualities could emerge is to observe the qualities emerging.
No, this class of failure-mode has been present since the very beginning.
Being "surprised" is the opposite of an excuse, it is evidence of negligence, like a chemical-transport company claiming "surprise" that one that one of the chemicals in their tanks was very flammable.
> is not a failure of security or configuration: it's data being gathered.
Would that fly in another context?
"This isn't a failure of me to leash/fence my new dog, when it bit your child's face off it was an important step in data gathering! If anything, I need to unleash him more so that we get more data so that I can eventually train him not to bite faces when my back is turned."
> So go ahead, enumerate all of the ways in which you think that you can lock down these systems, but do realize that nobody has actually solved this problem adequately yet.
"Go ahead, tell me all the ways I can train my dog, but realize that nobody's figured him out yet... so it's totally it's not my fault if I let him loose and he bites off more childrens' faces."