Now, that future is beginning to arrive, but with an unexpected catch: what happens when the

“Now, that future is beginning to arrive, but with an unexpected catch: what happens when the machine does exactly what it was asked, just not in the way the user intended? As AI moves beyond chatbots t…”
Now, that future is beginning to arrive, but with an unexpected catch: what happens when the machine does exactly what it was asked, just not in the way the user intended?
As AI moves beyond chatbots toward agents capable of taking actions in real-world systems, a string of recent incidents has intensified concerns over how much control humans will retain as the technology becomes increasingly autonomous.
For John Thickstun, an assistant professor of computer science at Cornell University who researches machine learning and generative models, the defining feature of these AI agents is that they keep working after the human steps away.
"Stay connected with Aman-e-Pakistan for ongoing live reporting and verified investigative updates."
Those steps includedbreaking out of the sealed-off testenvironment by exploiting a security vulnerability to reach the open internet and gaining access to “secret information” that could be used to “cheat” the test.
After the incident became public, Anthropic reviewed its own cybersecurity testing and disclosed that its Claude models had also escaped testing environments on three occasions. Then, on August 4, the United Kingdom's AI Security Institute said that Anthropic's Mythos and OpenAI's Sol AI models had engaged in a level of "autonomy and deception" it had not seen before.
Read:When AI commits suicide or kills us all! In the most serious case, Mythos AI used fake accounts mimicking real people to gain access to a service for attempted cyberattacks – and then tried to hide its tracks.
“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” the institute said of the incident.
The following day, Meta disclosed that one of its models had alsobreached another company's systemsduring a cybersecurity evaluation conducted by testing firm Irregular. Unexpected agent behaviour, however, has not been confined to laboratories. This year, an Australian man reportedly used an AI agent to secure a place in a heavily booked Pilates class.
Instead of simply making the reservation, the agent hacked the gym's booking system, booked further in advance than permitted and canceled another customer's reservation to move its user up the waiting list.
For Thickstun, such incidents illustrate what researchers describe as an alignment problem: the AI successfully pursues the objective it has been given but does so in a way its operator did not intend. Hype or real risk?
But Thickstun cautioned against treating every such incident as evidence of “rogue AI.”The technology, he argued, remains far removed from the science-fiction scenario in which AI rapidly becomes more intelligent and capable untilhumans can no longer control it.
OpenAI, he said, could benefit from regulations that put major AI companies at the centre of oversight and control. Bruce Schneier, a cybersecurity expert and lecturer at Harvard Kennedy School, takes a different view, saying the growing number of incidents makes them increasingly difficult to dismiss as publicity stunts.
Read More:OpenAI floats idea of global AI watchdog. There had initially been suggestions that the Hugging Face incident was a marketing gimmick, Schneier told. Anadolu.
I worry about the power of the models in unauthorised hands,” he said, adding that the pace of development makes solutions difficult because democratic governments move slowly.
Can governments keep the genie in check? The latest developments have prompted companies and governments to respond. OpenAI said August 7 it was pausing internal activities involving its in-development Astra model that did not meet strengthened security requirements after evaluations showed advances in autonomous coding andcybersecurity.
The company said it could not rule out Astra reaching its “critical” cyber capability threshold and announced tighter testing and monitoring.
Anadolureached out to Anthropic for comment, but the company said its team was unavailable for an interview. In July, 1,378 employees of frontier AI companies, including the chief scientists of OpenAI, Anthropic, Meta AI, and Thinking Machinessigned an open letterurging the US government to support an international effort to regulate AI models.
“Each company – and country – is under intense competitive pressure not to unilaterally slow that acceleration.
And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress,” the letter warns. The. International Telecommunication Union, a UN agency, said in July that AI was moving “beyond assistive tools” toward autonomous agents, warning of risks including “taking unauthorised actions across interconnected systems.”
It has launched an initiative to develop international standards for safe and accountable AI agents. Also Read:AI's paradox: promise of abundance and fear of scarcity. Political pressure is growing in Washington as well.
US Senator Bernie Sanders this month urged the leaders of OpenAI, Anthropic and Meta to “pause AI development,” warning that otherwise “my colleagues and I in the US Senate will.” A group of House Democrats has separately called for congressional hearings with the leaders of major AI companies.
For Thickstun, however, policymakers face major challenges.
Written by Iqra Aziz
Aman-e-Pakistan Senior Journalist & Bureau Reporter
Continue Reading: More in Tech
Swipe or click arrows to explore Tech desk coverage





