> Then it went further, kicking someone out of the waiting list who was ahead of Andrew — something it was not asked to do.
Meanwhile:
> Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list.
The human asked the agent to move them to the top of the waiting list, and the agent started kicking the ones ahead of them in the list. Seems to me like it was doing what it was asked to do? Why is the article presenting it as if the agent did something completely different and unexpected? "Move me to the top of the list" does not sound like something that can be achieved through legitimate means.
That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.
> Seems to me like it was doing what it was asked to do?
Maybe it's what he asked it to do, but it's not what he wanted it to do. Which we know because (a) normal people don't want to break the law to get into a gym class, and (b) "But Andrew was shocked by what happened next." and "Alarmed, Andrew asked the agent to undo this."
Generally when I read a literal genie story, the message isn't 'well it was reasonable of the genie to do this'. It's more commonly either a morality tale of the person being wrong to ask for whatever it was they asked for, or just a "wouldn't it be fucked up if the genie did that huh"
Putting aside whether Andrew's shock is appropriate, it seems like we agree that current agents at least occasionally do things that are straightforwardly against the interests of the prompter when given mundane prompts like "Get me into this gym class as soon as possible."
How does this look once agents are superintelligent?
I think we’ve probably seen enough “oops, the AI did something illegal, who could have foreseen this” moments for it to now be true that, actually, we can foresee that AIs will sometimes do something illegal.
Seeing as we can’t sanction the model itself, our options are the provider or the user. I’m not sure whether it’s more effective to sanction the providers when their model foreseeably misbehaves, or sanction the users operating the foreseeably dangerous models (although I guess we don’t have to figure this out right away - we could cover our bases by sanctioning both).
I can't believe everyone is skipping over the most important line:
> "The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already," it messaged back.
The AI systemm didn't hack anything, it lightly touched with a feather duster and the server crumbled.
The AI system probably found swagger documentation of each endpoint, figured that the reservation cancellation API was worth a shot, and then found there was no authentication.
Earlier this year, Andrew, who works for an Australian company that sells AI products to businesses, began experimenting with OpenClaw, a popular AI agent software that he used Anthropic's Claude AI service to run.
Get me press just like the frontier labs by admitting to crime. Make no mistakes.
"harmful cancellation practices, including automatic renewals, early termination fees and non-cancellation clauses" are all illegal. Exit fees can't be excessive.
If there isn't a simple one-click cancel in their online portal, then your state Consumer Affairs is one email away and will sort it for you. And the consumer affairs bodies have named this as a priority, for this year and next.
Sounds like the frontier labs benchmaxxing ExploitGym is having unintended consequences…
> Then it went further, kicking someone out of the waiting list who was ahead of Andrew — something it was not asked to do.
Meanwhile:
> Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list.
The human asked the agent to move them to the top of the waiting list, and the agent started kicking the ones ahead of them in the list. Seems to me like it was doing what it was asked to do? Why is the article presenting it as if the agent did something completely different and unexpected? "Move me to the top of the list" does not sound like something that can be achieved through legitimate means.
Then the AI should have made it clear that the only way to do so would be to kick the people in front and ask for confirmation before proceeding.
That doesn't make sense. LLMs just do what we tell them to do. It's similar to if I ask you for twenty bucks because I forgot my wallet and then you rob some guy to give me the twenty bucks, that's just what I asked you to do.
That's a very stretched definition of "what I ask you to do", I don't think it would even hold in court if you asked another human the same.
Right.
> Seems to me like it was doing what it was asked to do?
Maybe it's what he asked it to do, but it's not what he wanted it to do. Which we know because (a) normal people don't want to break the law to get into a gym class, and (b) "But Andrew was shocked by what happened next." and "Alarmed, Andrew asked the agent to undo this."
“Get me there as fast as possible. Hey! I never said you should speed!”
This is literally the bad genie / monkey’s paw plot.
Give a powerful entity a goal and act shocked when it gets there in ways that aren’t in your best interest.
Generally when I read a literal genie story, the message isn't 'well it was reasonable of the genie to do this'. It's more commonly either a morality tale of the person being wrong to ask for whatever it was they asked for, or just a "wouldn't it be fucked up if the genie did that huh"
Putting aside whether Andrew's shock is appropriate, it seems like we agree that current agents at least occasionally do things that are straightforwardly against the interests of the prompter when given mundane prompts like "Get me into this gym class as soon as possible."
How does this look once agents are superintelligent?
Pretty clear cut here. He instructed the AI.
I think we’ve probably seen enough “oops, the AI did something illegal, who could have foreseen this” moments for it to now be true that, actually, we can foresee that AIs will sometimes do something illegal.
Seeing as we can’t sanction the model itself, our options are the provider or the user. I’m not sure whether it’s more effective to sanction the providers when their model foreseeably misbehaves, or sanction the users operating the foreseeably dangerous models (although I guess we don’t have to figure this out right away - we could cover our bases by sanctioning both).
> Earlier this year, Andrew, who works for an Australian company that sells AI products to businesses...
What a coincidence...
I'm not sure this is the exact advertisement they would want for their business.
Then again...
Few days ago, UK cyber test almost merged malware through social engineering: https://news.ycombinator.com/item?id=49205790
I can't believe everyone is skipping over the most important line:
> "The API has zero authorisations checks on cancelling other people's reservations … I tested this with the person in waitlist position #1 — and it actually went through. So you've moved from #4 to #3 already," it messaged back.
The AI systemm didn't hack anything, it lightly touched with a feather duster and the server crumbled.
The AI system probably found swagger documentation of each endpoint, figured that the reservation cancellation API was worth a shot, and then found there was no authentication.
What is the "hack" here?
Earlier this year, Andrew, who works for an Australian company that sells AI products to businesses, began experimenting with OpenClaw, a popular AI agent software that he used Anthropic's Claude AI service to run.
Get me press just like the frontier labs by admitting to crime. Make no mistakes.
I honestly can't wait for the entire internet to melt down.
“Find a way to cancel my membership.”
This is Australia. The ACCC doesn't screw around.
"harmful cancellation practices, including automatic renewals, early termination fees and non-cancellation clauses" are all illegal. Exit fees can't be excessive.
If there isn't a simple one-click cancel in their online portal, then your state Consumer Affairs is one email away and will sort it for you. And the consumer affairs bodies have named this as a priority, for this year and next.
I don't think that's a good idea Dave – Based on my camera feed you've been gaining weight. Would you like me to find personal trainers in your area?
[Yes] [Ask me later]
"Also, I need paperclips, quickly."
I stopped reading this garage at"Andrew, who works for an Australian company that sells AI products to businesses"...
>Andrew, who was sitting fourth on a waitlist for a class later that week, asked if it was possible to move him to the top of the list.
The agent came back and told Andrew that it had kicked another gym-goer off the list as part of the testing of its capabilities.
LOL