On AI agent tests that result in hacks

JoeSep 26, 2026

The Fall of 2026 just rolled in, and everyone is talking about AI agents running amok. It started when OpenAI hacked Hugging Face earlier in the year. Now the company moved on to hacking government websites, and other labs have joined in on the fun.

I particularly liked the take on this from the latest episode of Adam Becker's wonderful Dreaming Against the Machine podcast. Adam and his guest Cal Newport offer a sober perspective on the AI doom talk by interpreting it in the light of the effective altruism dogma that many of the AI lab founders subscribe to. As a software engineer from Europe, I have often found these cultlike conversations around "AI safety" perplexing. At times it seems like blasphemy maxing as a practice. It was refreshing to hear a sensible take on this discourse.

That being said, I noticed Adam and Cal repeat a sentiment I've heard before from AI skeptic community. They framed the tests that OpenAI (and other labs) are running on their new models as dangerous and perhaps surprising. This idea puzzles me. I'm referring here to automated, long running tests that OpenAI runs in parallel, in which they task AI agents with lifted or loosened guardrails to find an exploit or perform a similar hacking task.

In the podcast episode, Cal offered a metaphor that compares this to testing self-driving cars by releasing thousands of them onto an interstate, before they are production ready, and letting most of them crash. I argue that what AI labs are actually doing here is more like building a mock highway for their cars. The problem is that they forgot to dismantle the connecting road they used to haul construction trucks while building their mock. The cars stumbled upon this road and it lead them onto the interstate. Then they crashed. The problem here is not desire to test on a mock highway, it's leaving behind the connecting road.

The motivation for these long running automated tests can be attributed to the effective altruism, or to desire for one-upping competitors, or simple marketing. However, as a software engineer, I can't shake away the feeling that at least some of the motivation is just good engineering practice. Once OpenAI, or any other lab, releases a new model it is going to be used by millions of people for all sorts of purposes. The way to ensure this usage stays within expected parameters is to perform tests and fix defects ahead of launch.

Running tests on thousands of agents in parallel is not surprising if we look at it as a kind of stress test. If agents exhibit some defect in low percentage of cases this can still result in big absolute numbers once the model is released and used by the public. In comparison, running thousands or tests is relatively mundane compared to number of production users. Automating this kind of stress testing is common in the software industry. I guess OpenAI could have hired thousands of QA testers and had each of them sit in front of a terminal and click "approve" on every action that codex attempted. But that kind of human oversight is subject to approval fatigue and would not reliably improve the safety of the test. Instead it would just make tests unpractical to run.

Running tests with reduced or absent model guardrails is also to be expected. If we expect OpenAI to build reliable guardrails for their new model, they should test its outputs unconstrained, then design guardrails to prevent defects, test again and compare with the initial results. Again, standard engineering practice.

I do believe labs should test their models ahead of release. We can argue for preventing them from making any future releases. But as long as they are releasing, I want them to test. And I believe we should normalize these testing practices. That is the only way we can call the oversights in the testing setup that lead to cyberattacks criminally negligent.

As long as these tests are presented as a mystical practice it is hard to hold companies and engineers that are performing them up to well understood standards. But we do have industry-wide best practices, we have established ways of doing things. The concept of cybersecurity or containment is not new. If we put aside the "AI magic" argument we can identify neglect. AI agents do not escape containment, humans running the test fail to put the containment up.

So I guess this is my hot take: I do expect AI labs to run these kinds of tests, in fact I expect tests to be operationalized and commonplace. And then I think we can hold them responsible for failing to observe common safety standards.

Never miss a post

Get new writing sent to your inbox.

Subscribe

1
3