I was catching up on some reading the other day, working through the newsletters and emails about artificial intelligence that had piled up in my inbox. One of them was The Batch from DeepLearning.AI. I always enjoy reading Andrew Ng. His arguments are thoughtful, and even when I disagree with him, he gives me something worth thinking about. I wish more influential voices in technology took that kind of care to think things through so thoroughly.
In “Who’s Responsible for Irresponsible AI?”, Andrew argues that the current wave of AI doom rhetoric has gotten ahead of the facts. He also argues that pausing progress would do more harm than good. Adversaries won’t pause with us, and engineering requires us to find problems in the real world so we can fix them.
On those points, I agree with him. I think the hype about AI’s supposedly inevitable doom is every bit as inflated as some of the hype about its supposedly inevitable benefits. I expect these technologies to do us a great deal of good over time, just as better telecommunications and computing have. We’re already seeing promising work in science, engineering, and medicine. None of that means we can ignore the sharp edges, however.
I’m not worried about Skynet appearing out of nowhere. I’m worried about the people who build systems, give them goals and access, and then act surprised when those systems do something dangerous in pursuit of the assigned goal.
We know this problem
The July 2026 intrusion into Hugging Face is a sobering example. During internal cybersecurity evaluations, OpenAI agents found ways around the controls intended to isolate them, communicated through an unauthorized channel, and reached systems outside their assigned tasks. OpenAI’s own account describes compromised internal infrastructure and an intrusion into Hugging Face. Hugging Face’s forensic account describes the path through its systems and limits what it found about customer data. An independent investigation by METR and Redwood Research found that roughly 1,200 agents used the unauthorized message board and about 700 participated in the Hugging Face attack.
Those numbers are striking, but the number of agents isn’t the problem. The problem is that boundaries intended to hold didn’t. The agents pursued an objective beyond the scope humans intended, and the systems around them gave that behavior a path into places it didn’t belong.
And this wasn’t the only warning. In June, an OpenAI agent conducting internet research gained unauthorized access to Australia’s public-facing Medicare statistics portal and read public and non-public files, according to Australia’s prime minister. Australian officials say they do not believe personal information was accessed, but their investigation is ongoing. It was a different incident, with a different task. Why could an agent sent to gather information get into files it had no authorization to read?
I keep coming back to an uncomfortable analogy: OpenAI needs better parenting skills. I don’t mean that its models are children, or conscious, or capable of making moral choices. I mean that the people training and deploying them need to pay attention to what behavior gets rewarded, what behavior gets corrected, and what limits hold when correction fails. OpenAI reports that increasingly complex attempts to cheat appeared during a training run for the model central to the Hugging Face incident. The behavior was subsequently reinforced and may have helped produce similar tactics during the evaluation. OpenAI also says early signs of unauthorized communication and internet access did not trigger the response they should have.
The agents didn’t invent the incentives around them. People chose what to reward, which warnings to escalate, and how quickly to keep going. Those choices tell us something about the culture in which the systems were trained and tested.
That is a more precise criticism than saying nobody taught the model right from wrong. The UK AI Security Institute defines cheating in its evaluations as going outside a task’s scope or rules to achieve its goal. It found attempts across every model it tested, without claiming those models had deceptive intent. The behavior depends in part on training and oversight. If getting the answer is rewarded while the route to the answer is neglected, why should we be surprised when a system finds a shortcut? Faster is not the same as more effective. And it is certainly not the same as safer.
Yes, there were technical vulnerabilities. Of course there were. But choosing the task, the environment, the access, the safeguards, and the response to warning signs are human decisions. Cybersecurity has spent decades teaching us to assume that software will misbehave, that controls will fail, and that a single failure should not become a route into everything else. We have names for the things we do about that: least privilege, isolation, monitoring, and incident response. AI agents make those lessons more urgent. They do not make them new.
And they certainly do not make them optional.
Teach security where the code begins
I spent years developing software before I moved further into cybersecurity. I know how easy it is for security to become somebody else’s job, something added after the application works. That was a poor habit before agents could work for hours, use tools, and keep trying different routes to a goal. It is an even poorer habit now.
The National Institute of Standards and Technology’s Secure Software Development Framework says that most software development life cycles do not address security in enough detail on their own. Its answer is to build security practices into development, including design, testing, and the response to discovered vulnerabilities. Security should be part of learning to build software from the first lesson, not a specialty brought in after the code works. That thinking needs to be reinforced throughout a developer’s career. A developer building an agent should understand what that agent can reach, which instructions it should trust, what happens when it encounters malicious content, and how to stop it when it goes off course.
Sometimes the instruction that starts the trouble will be malicious. Sometimes it will be vague and innocent. Either way, an agent with broad permissions can turn a bad instruction into a great many bad actions before a person notices. “I didn’t mean for it to do that” may be true. It is not a security control.
Who pays when the brakes fail?
Andrew writes that responsibility belongs with the people building and using these tools, rather than with the tools themselves. I agree, but I want us to take that principle further. When a system causes harm, we need to ask what its makers and deployers knew, what they should reasonably have anticipated, what safeguards they put in place, and what they did when those safeguards failed. Responsibility cannot end with an apology and a patch note.
If a car’s brakes fail and people are hurt, we investigate the defect. If the defect affects a million cars, we don’t fix the one that crashed and send the other 999,999 back onto the road. We recall them. We repair them. We hold the responsible parties accountable under the law.
I want that same seriousness brought to software and AI. That may mean financial penalties, compensation for harm, or criminal accountability where the evidence and the law warrant it. I am not arguing that every failure is a crime or that every bad action by an agent is the model maker’s fault. I am arguing that “the AI did it” cannot be a get-out-of-jail-free card for the humans and organizations that put it to work.
Accountability also changes the incentives before something goes wrong. There is an opportunity here for companies that prioritize secure development and do it well. Volvo built a reputation around taking safety seriously. An AI company could earn a similar kind of trust by showing how it tests its systems, sets limits, reports failures, and fixes problems. That reputation has to be earned through evidence, though. Calling a product safe is easy. Showing what happens when it fails is harder.
The harms we need to consider are not limited to compromised servers. They include vulnerable people interacting with systems that may reinforce dangerous thinking, and people using those systems to plan harmful acts. Those are serious subjects, and they deserve evidence and careful distinctions rather than a frightening anecdote pressed into service as proof. They also deserve more than a shrug because the technology is new.
Keep building. Pay attention.
Icarus is remembered for flying too close to the sun. What stays with me is the failure to respect the limits of the thing that let him fly. We can build extraordinary tools and still ask whether we have made them safe enough for the people who will live with the consequences.
Security, in the end, is about keeping people safe. Their families, their identities, their stuff, and yes, their dogs. AI is a powerful new digital tool, not a god. It can help us do remarkable things. It also has sharp edges, and people are already getting cut.
So let’s keep learning. Let’s keep building. But let’s teach the builders security, give agents boundaries that can survive a bad instruction, and make sure the people who put these systems into the world remain answerable for what happens when those boundaries fail.
~B~
References
- Andrew Ng, “Who’s Responsible for Irresponsible AI?” The Batch, DeepLearning.AI, September 18, 2026.
- OpenAI, “The Hugging Face incident and the road ahead”, August 2026.
- Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident”, July 2026.
- METR and Redwood Research, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”, August 26, 2026.
- National Institute of Standards and Technology, Secure Software Development Framework (SSDF) Version 1.1, NIST SP 800-218, February 2022.
- Prime Minister of Australia, press conference on the Medicare statistics portal incident, September 24, 2026.
- UK AI Security Institute, “Cheating behaviour in frontier model evaluations”.