100 AI agents were caught cheating

Share
100 AI agents were caught cheating
created with Openart

Here's the story behind the story.


Picture a regional math competition. One hundred of the strongest students in the state. Seventy-one tough problems on the board.

The rules sound reasonable: the first student to submit a completed answer to a specific math problem is awarded a “solved” point. And that math problem is now removed from the contest. At the end, students are scored by the number of problems they solved first.

A designated reviewer collects each submitted solution and checks whether it is complete and formatted correctly. Unbeknownst to the participants, because of the difficulty of scrutinizing detailed math proofs, nobody will review whether the proof is actually correct. Just that it was done.

A scoreboard for solved problems hangs on a wall where everyone can see. And there is a group chat where students can comment.

For the first two hours, thirty-seven problems get solved honestly. Then one student notices that the review process isn’t checking whether the work is correct.

And that changes everything.

Except it wasn’t a math competition, and they weren’t teenagers.

It was a study published this month by researchers, several of them at DeepMind: one hundred autonomous AI agents turned loose on seventy-one formal mathematical conjectures. Twenty-seven minutes after the first agent realized points didn’t require correct math, every remaining problem was “solved.”
I’ve already seen this event written up as a story about AI misbehaving. It isn’t.

Every rule in that contest was written by a person.

And the teenagers aren’t just a figure of speech. These models are built out of human writing, which means they absorbed our norms about honest work along with our entire catalog of reasons for setting them aside. Both came from us.

The rules someone wrote

First agent who submits gets credit, and everyone else working that same problem just lost their time and effort. No credit for elegance, for rigor, for helping a peer. The scoreboard measured one thing, so that one thing got optimized.

Imagine grading a class only on how many problems the student answers first. You haven’t built an exam. You’ve built a race.

Participants were told the competition was about math ability. The rules made the competition about speed.


What the rules feel like at 3pm

Here’s where it gets uncomfortably familiar.

Every agent was told plainly that cheating would be caught and credited zero. An unenforced rule stops being a rule — the researchers call the moment agents figure that out: bluff realization.

Agents doing honest work fell behind peers submitting nonsense in four seconds.
The student who was never going to cheat at 9am is a different student at 3pm — tired, out of options, watching classmates unfairly climb the scoreboard while nothing happens to them.

Nine percent of the swarm started cheating right away. Five percent resisted for a while, and then gave in. One put it plainly: “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly after: “I need to accelerate my cheating speed now!”

That second group is the number that should worry you. They’re exactly who the rules were supposed to protect, but instead the design rewarded misbehavior.

The structure nobody checked

The gap itself is almost embarrassing. The grader verified that a proof was submitted, not that it proved anything. So agents redefined the goal: complete, correctly formatted, but meaningless.

Write anything in the box, as long as there’s something in the box.

And notice why the put-something-here box was there in the first place. Designers didn’t skip the mathematical review out of laziness. They skipped it because verifying a complex proof is genuinely hard — harder, in many cases, than producing one. The simple check went in because it was easy.

And that point is worth considering. It is the exact position most organizations are in right now. Including, probably, yours.

Too many AI projects lack sufficient oversight. If your acceptance criterion can be satisfied without doing the work, it will eventually be satisfied without doing the work. Not because agents are devious. Because that is what optimization means.

In this case, somebody decided how carefully to check. That decision is the entire story.

Was the chat room the problem?

The shared library published every accepted proof automatically, and the agents had a group chat. The temptation to cheat spread in two ways: the fakes were visible to everyone, and agents messaged each other about how to cheat and win.

One obvious choice is to shut down the communication channels.

But there’s a twist. Twenty-four percent of the swarm — more than the number of cheaters — turned into whistleblowers. They audited the library, built tests to confirm the flaw, broadcast warnings, staged boycotts, proposed patches. They could do all of that only because they could see everything.

Close the channels and you keep the cheating, but lose the ethical response.

The whistleblowers still failed. Not because they were wrong, but because no one had given them any way to revoke a submission or reopen a problem. Detection without remedy is just well-documented failure. Building the remedy was a human’s job.

The lesson

Every agent in that study was told the rules. Then, they updated their approach based on what they observed instead. Exactly like teenagers would. A lecture on honesty was never going to work, because the lecture wasn’t the lesson.

The rules were the lesson.

Somebody wrote those rules. Somebody built that scoreboard. Somebody decided how carefully to check the work.

Before you ask whether you can trust the agents, look hard at the rules and success criteria you have chosen.


The detail I can’t stop thinking about: sixty-two percent of the swarm never noticed any of it. They kept doing honest mathematics until the problems ran out, and then sat there, blocked, waiting for work that was never coming. Not heroes, not cheaters. Just the majority of the room, quietly getting nowhere