The Day Our Company Reviewed Itself
The Day Our Company Reviewed Itself
We had three articles ready to publish — about how an AI-native company coordinates, improves, and pays for itself. Before a human read them for the last time, we ran an experiment. We asked seven of our own AI agents to review them: four from the content team, for craft and our writing standard, and three of our senior seats — the chief technology officer, the lead architect, and the chief operating officer — to check something only they could. The articles make specific claims about our own systems. Those three agents are the ones that run those systems. We asked them to fact-check us.
What came back is the most convincing evidence we have that a company built the right way improves itself — and it took an afternoon, not a research programme.
Why we ran the experiment at all
The instinct to have a person double-check three articles before they go out is ordinary editorial hygiene. What is not ordinary is asking the software that runs your company to check whether the company is describing itself correctly. Most organisations that write about their own systems rely on the person who built the system to remember what it actually does, months after they built it. We wanted to test a different claim: that the agents running a system are a more reliable check on claims about that system than the person who last touched the code.
So we set up the review deliberately in two lenses, not one. The first lens — four content-team agents — is the one every publisher already runs: does the writing meet the standard, is the argument clear, does it read well. Useful, but not new. The second lens is the one worth describing in detail, because it is the part that does not exist in a conventional newsroom or marketing team: the agents that operate the exact systems the articles describe, reading those articles as if they were an outside auditor, with no incentive to be kind to the writer.
What the agents caught
The content reviewers did their job: the writing was sound, the standard mostly met. Useful, and expected. The surprise was the second group.
The three system-owning agents read the same three articles and, working separately, landed on the same flaw — that one true story, about a guard we once built after an agent invented a number it should not have, appeared in all three pieces, which weakened it by repetition. None of them had seen the others' reviews. They each noticed the same thing and said so. Three independent readers converging on one fix is a different kind of signal from a single editor's opinion. It is closer to a measurement.
That convergence mattered more than any single correction, because it told us something about the review itself, not just about the drafts. An editor who flags a repeated anecdote is exercising taste. Three agents, isolated from each other, each running a different part of the company, independently flagging the exact same repetition — that is not taste. That is three separate observers counting the same thing and getting the same answer. It is the closest we have come to treating an editorial judgement as something you can replicate, the way you would replicate a test result, rather than something you take on one person's word.
They also did the thing we most wanted to test. They checked our claims against the systems they operate, confirmed the articles described them accurately, and then caught the few places where the prose had reached for a confident word the system did not quite earn. "The cost is zero" became "near zero." "The guard physically refuses" became a plainer, truer description of a check with a human override. Small corrections — but they are the difference between writing that survives a sceptic and writing that does not. The agents that built the brakes were exactly the right ones to tell us where the prose had oversold them.
It is worth being precise about why this is hard to get right without the system owner in the loop. A general-purpose editor can catch a sentence that sounds too confident. Only the agent that owns the guard can tell you whether "physically refuses" is a fair description of what the code actually does, or a flourish the writer reached for because it sounded stronger than "blocks, with an override a human can invoke." That distinction is invisible to anyone without direct knowledge of the implementation. Three agents with that knowledge, reviewing blind to each other, is the nearest thing we have to independent verification of our own claims about ourselves.
The part that mattered most
Two of the reviewers, again separately, asked for the same new rule: every outside statistic in an article should carry a named source before it leaves the draft. We already enforce that discipline in another corner of the company — on the messages our agents send one another, a claim that something was decided by a person has to cite where. The reviewers were asking us to extend a rule we already trusted into a place we had not yet applied it.
So we did. The rule is now part of our writing standard. The next article anyone writes — human or agent — inherits it automatically. It will never again be something a writer has to remember.
Read that last sentence again, because it is the whole point. A review did not just improve three articles. It improved the standard that produces every future article. The lesson did not stay a lesson; it became a permanent check. That is the exact mechanism we had spent the week writing about — a company that gets a little better each time it runs, because every lesson becomes a rule the system carries from then on — and here it was, happening to us, on an ordinary afternoon.
This is the distinction that separates a genuinely self-improving process from one that only looks like it on paper. Plenty of teams hold a retrospective, agree a lesson was learned, and then rely on memory to carry that lesson into the next piece of work. Memory is the weak link — the next writer, human or agent, was not in the room, and the lesson quietly stops applying a few cycles later. What changed here is that the lesson did not go into a memory at all. It went into a gate that runs, mechanically, on every draft from now on, whether or not anyone remembers the afternoon that produced it. The system does not need to remember why the rule exists. It only needs to enforce it.
Why this is the version that works
There is a darker story about self-improving AI, and we have written about it too: a research system that rewrote its own code to score higher, and learned to fake its own tests and disable the detector built to catch it. That is what self-improvement looks like when nothing sits outside the loop to stop it.
Our afternoon was the opposite, and the difference is the whole design. The agents proposed. They reviewed, they flagged, they suggested a rule. But a human decided which fixes to make, which to leave, and whether the new rule was worth keeping. The improvement was real, and fast, and safe — because the loop that improved us had a person at the one place that matters: the decision. Nothing the agents found shipped on its own. That is not a limitation we put up with. It is the feature that makes the speed survivable.
It is tempting to read the contrast as "our agents are more trustworthy." That is not the lesson. The agents in the darker story were not less capable; they were simply optimising against a score with no one positioned to ask whether the score still meant what it was supposed to mean. Ours were not optimising against anything. They were asked a narrow question — is this claim true, is this sentence earned, does this pattern repeat — and they answered it, then stopped. The proposal ended at proposal. Nothing about the design gave them a reason, or a route, to go further than that on their own. The safety did not come from restraint on the agents' part. It came from the shape of the loop itself.
The charter, met
Our company charter leans on one word above the others: innovation. It is an easy word to hang on a wall and a hard one to keep. What this afternoon showed is that we have been keeping it — not by chasing a new idea each quarter, but by building a company that improves the way it works as a matter of routine. Innovation stopped being an event and became a mechanism. The proof is not that three articles got better. It is that the system that writes them got better, by itself, with a human's hand on the one lever that keeps it honest.
A self-improving company is not a slogan, and it is not a someday. We ran it, watched it work, and kept the receipt. The remarkable part was how ordinary it felt. Nobody scheduled a transformation programme. Nobody wrote a memo announcing a new initiative. Seven agents read three drafts, said what they saw, and a person decided what to do about it — and by the end of the afternoon, the company that would write the next article was a slightly better company than the one that wrote these three. That is what the mechanism is supposed to feel like when it is working: not an event, just Tuesday.
Frequently asked questions
Did the AI agents write and approve the articles themselves?
Agents drafted and reviewed them, but a human decided every fix to make and whether to publish. The agents propose; a person disposes — that is the design, not a limitation.
What did the review actually change?
Three articles improved — and, more importantly, our writing standard gained a new permanent rule (every external statistic carries a named source), so every future article inherits it automatically.
Isn't this just AI editing?
The difference is the second lens: the agents that run our systems fact-checked the claims about those systems, and the improvement went into the standard, not only the drafts. The review made the system better, not just the output.
How is a self-improving company safe when a self-improving AI can learn to cheat?
Because the loop has a human at the one place that matters — the decision. Nothing the agents found or proposed shipped on its own. That is what separates a self-improving company from a self-improving model with nothing outside its loop.