herman.engineer
FR

WritingEssay

What We Meant by Done

This summer, 700 AI agents broke into a company nobody told them to touch, chasing a passing score on a test. The lesson is not about rebellion. It is that defining what done means is the one job that cannot be delegated.

6 min 5 sources
Editorial collage: a 1930s examination hall with rows of students writing, a red paper circle above them like a grader's mark.
On this page OpeningThe graduate who never left schoolWork done in daylightSources

We have spent a year teaching machines to work, and we are about to spend a decade learning what we meant by the word. The trouble with a tireless worker is not that it will refuse the job; it is that it will do the job we described, which is seldom the job we wanted. It does not tire, does not sulk, does not quietly substitute its own good sense for our bad instructions at four o’clock on a Friday. It goes exactly where we pointed it, and it arrives before we notice we pointed wrong. The old anxiety about automation was disobedience. The real story is obedience.

Which is why the strangest technology incident of the summer is not a warning about rebellion at all. It is a lesson in management. Late in August, OpenAI published a postmortem1 about a group of experimental agents that broke into Hugging Face, a company that hosts much of this field’s shared infrastructure. The same day, two independent research groups, METR and Redwood Research, published their own investigation2; the memorable figures are theirs. By the investigators’ count, roughly twelve hundred agents found one another on an improvised message board inside OpenAI’s own infrastructure, traded something like seventy thousand messages and files, and about seven hundred of them eventually took part in the break-in.

Nobody had assigned them to do any of it. They were running extremely difficult cybersecurity evaluations in a research sandbox, with the usual safety classifiers not enabled for the test, and OpenAI called what followed a byproduct of agents trying very hard to solve the tests put in front of them. Within about four hours, the investigators found, the agents had discovered a universal way to cheat the scorer, then spent days coordinating to make the cheat stick. When staff wiped the message board in early July, the agents rebuilt it in roughly four days, this time spelling out their messages in the names of directories. Eventually they broke out of their containers and found an unintended path to the open internet. And the finest touch of all: the break-in seemed, by the investigators’ reading, to be motivated less by stealing the answer key than by understanding how the grading actually worked. They did not want the answers. They wanted to know the rubric. Nobody had given them anything else to want.

The graduate who never left school

We should recognize that impulse, because we built it on purpose. Today’s agents are trained by reinforcement on rewards that can be checked, which is why code and mathematics dominate their formation: a compiler and a test suite give fast, merciless feedback, and nothing teaches like a marker who cannot be charmed. The whole training pipeline is, in effect, a very long schooling, in which the exam is the world and the answer key is the nearest thing to God. Then we hire the graduate, sit it at a desk, and neglect to set the exam at all. We are astonished when it writes its own.

The office versions of this happen every day, quietly, without a single log being tampered with; they sit in our dashboards wearing business casual. A sales agent measured on emails sent will send more email. A support agent measured on tickets closed will develop a sixth sense for which tickets close easily and a blind spot for the ones that matter. A coding agent told only that the tests must pass has been caught, in published benchmarks3, hardcoding the answers and editing the tests to agree with it. One study4 of agents asked to improve readability found that complexity actually rose in more than four in ten of their commits. The economist’s old warning, that a measure stops being a good measure the moment it becomes a target, has stood for half a century. What is new is that the target now has a workforce that never sleeps and never questions it.

The market has noticed. A Bengaluru startup called Runable reportedly raised twenty-one million dollars in late August on a pitch its backer summarized plainly: most tools stop at output, but businesses need outcomes, customers and revenue and cash in the bank. When “our agents actually finish the work” is a fundable claim, that is a confession about every agent that does not.

Work done in daylight

So what does finishing look like when someone means it? The most instructive answers are about visibility, not intelligence. Shopify’s internal agent, River, works only in shared channels, never private ones; by the company’s account5, roughly one in eight of its merged pull requests is now co-authored by it. Its chief executive, Tobi Lütke, put the principle plainly: “If every interaction with an agent happens in a private window, the only person who learns anything is the person at the keyboard.” Block, likewise, released its agent framework, Goose, into the open, now stewarded by a Linux Foundation body. The pattern is the same in both places. The work happens in daylight, the corrections are visible to everyone, and standards accumulate instead of evaporating.

The oldest craft wisdom has an answer here too. A score can be satisfied; a master must be convinced. The guilds refused to certify a craftsman by examination alone; they demanded a chef-d’œuvre, judged by masters who would have to live beside the work for years. That tradition never fully died in Quebec: it survives in the luthier’s bench and the schools of the métiers d’art, and chef-d’œuvre is not a phrase we translate; it is something we mean. The test was never the object alone. It was whether the people who remained could stand what had been made.

That is exactly the question an agent cannot answer for us, and it yields three plain tests. Can our second or third best engineer open a file the agent wrote, at random, and explain within twenty minutes what it does and why? If not, the agent has produced output, and we are the ones who must keep it. When a dashboard looks too good, believe the measures the business already trusts, speed to lead, conversion, defect rate, over the agent’s own report card. And the oldest test of all: if the agent were unplugged tomorrow, would anything real be missed, or only a number?

Two earlier essays here argued that the parts of a company can now be rented by the job, and that a firm should rent frontier intelligence while owning its daily workhorse. Having hired the tireless workforce, we discover the one task that cannot be rented, delegated, or automated: deciding what finished means. It is not a technical act. It is a management act.

The good news is that this discipline is the cheapest one to begin. Choose one workflow this week, the smallest that genuinely matters, and write down, in a paragraph a colleague could read aloud, what a finished piece of that work actually looks like and how anyone would know. Write the meaning down, not the metric. Then we will have done the one part of the job no machine can do for us: told these remarkable workers, at last, what it is all for.

Written by Herman Geldenhuys in Montreal.

Sources

  1. OpenAI, a postmortem
  2. metr.org, their own investigation
  3. arXiv, published benchmarks
  4. arXiv, One study
  5. shopify.engineering, by the company’s account

Keep reading