Meta wanted AI agents to do the work of thousands. It did not go to plan
Reuters reported that Meta shelved its AI agent restructuring. Its own internal numbers explain why, and they are worth understanding.
Two numbers turned up next to each other in Meta's internal reporting this year. Code changes to the company's own platforms and infrastructure were up 220 percent year on year. Changes that actually reached users as new or improved features were up 36 percent.
That gap is the whole story of a Meta project code named OT, and it is one of the most useful things published this year about what AI agents currently do to a large organisation. It arrived without a press release or a keynote, from a company that had no interest in releasing it.
Where this account comes from
Reuters broke the story on 26 August 2026, drawing on internal material and on interviews with people who had knowledge of the project. We could not load the Reuters page directly, so the details here reach us through the outlets that carried the reporting on and after that date, including Engadget, TechSpot and The Next Web, all of which credit Reuters.
That matters for how you should read what follows. These are internal figures described by reporters who saw internal material. Meta has not published them, has not confirmed them in a statement we could find, and did not choose to release them. Everything below is reported, not announced.
A retreat in Hawaii
Project OT stood for Organization Transformation. According to the reporting, it took shape at a leadership retreat at Mark Zuckerberg's Hawaii estate in January 2026, where executives laid out an "AI native" vision for how the company could be rebuilt around software agents.
The shape of it was specific enough to be startling. Product teams of ten to twenty specialists would collapse into pods of three to five people. The job titles that structure a modern tech company, engineer and designer and product manager, would consolidate into a single role called "builder". Layers of middle management would be removed on the theory that agents could absorb the coordination work those layers exist to do. Daily priorities would be set by what the internal material called agent assisted analysis.
Read charitably, it is an honest attempt to take AI seriously. If agents can genuinely do a large fraction of the work, then the org chart built around humans doing that work is the wrong shape, and pretending otherwise is its own kind of denial. Meta decided to find out.
Two waves, and the one that never happened
The plan ran in two stages. The first went ahead: roughly 8,000 roles cut, about 10 percent of the workforce, with around 7,000 people reported to have been moved into AI focused roles in the same week. That part is real, and the people affected by it were affected by it permanently.
The second stage was scheduled for November. It was considerably more aggressive, reportedly cutting some teams by as much as 60 percent through a mix of layoffs, hiring freezes and the removal of people the company rated as poor performers. It never ran. The Next Web's account of the Reuters reporting has Zuckerberg calling it off on the night of 19 May, hours before the first wave landed the next morning. Engadget's own contemporaneous report of the 10 percent reduction carries an April date instead, so treat the calendar as unsettled.
The sequence holds either way. The November stage was killed before the first one had finished landing, quietly, and against a plan the company had spent months building. Something in the internal data had made the case for stopping.
220 against 36
Here is what that something appears to have been.
Code changes to Meta's internal platforms and infrastructure rose 220 percent year on year. Changes that reached users as new or upgraded features rose 36 percent. Major technical and security incidents rose 40 percent. Time spent by employees dealing with those incidents rose 70 percent. The Next Web traces the 220 percent figure to a June internal post by Meta's chief technology officer, Andrew Bosworth. The other three appear in the reporting without a named author.
Line those four up for a second. Activity more than tripled. Delivery went up by roughly a third. And the cost of the extra activity shows up plainly in the last two figures, because a 40 percent rise in incidents with a 70 percent rise in time spent cleaning them up means the incidents were also getting harder to clean up.
Why volume and value came apart
The gap has a fairly ordinary explanation, and it generalises well beyond software.
Agents are extremely good at producing units of work. A code change is a unit of work. It is discrete, it has a clear beginning and end, it can be generated at volume, and it is easy to count. This is precisely the kind of thing a capable model does well, and a 220 percent increase is genuinely impressive as a measure of throughput.
A shipped feature is a different object. Getting there requires deciding whether the thing is correct, whether it is safe, whether anyone wanted it, whether it breaks the system sitting next to it, and whether the company is willing to put its name on it. Those are judgements, and judgements bottleneck on review, integration and someone being accountable for the outcome. Pouring more input into the front of that pipeline does not widen the part of it that was already narrow. It just makes the queue longer, and the incident figures suggest that some of the queue got processed less carefully than it used to be.
The human side of the story points the same way. Meta's half year Pulse survey reportedly fell from 74 percent favourable sentiment to 55 percent, with staff objecting in particular to tracking software mandated on employee devices to capture keystrokes and mouse clicks, so that agents could learn to use a computer the way a person does. People who believe they are being measured in order to be replaced tend to stop volunteering the judgement that was holding the pipeline together.
Zuckerberg himself said something notably undramatic about it at a July town hall:
The trajectory of the agentic development over at least the last four months hasn't really accelerated in the way that we expected.
Neither a victory lap nor a funeral
It is worth being careful about what this episode proves.
It does not show that AI agents do not work. Meta did not reverse the May cuts, did not stop moving people into AI roles, and has kept shipping AI products through the summer, including an open weight model called Muse Glimmer released on 10 August that is built to run on consumer hardware. A company that had concluded agents were useless would behave differently.
It also does not show that the technology quietly transforms everything while sceptics look the other way. One of the best resourced engineering organisations on earth restructured around agents, measured the result honestly, found that the line it cared about had barely moved, and stopped.
And it says nothing at all about anyone else. Reuters reported on Meta. Whether other companies see the same pattern is a question nobody has published comparable internal numbers to answer, and filling that silence with confident extrapolation is how the hype cycle got here in the first place.
What to take from it if someone is telling you AI will change your job
The transferable lesson is about measurement, and it is the part most likely to show up in your own working life.
If an organisation decides to measure AI by volume of work produced, that number will go up. It is almost guaranteed to go up, because generating volume is the thing these systems are best at. The number will be reported in an all hands meeting, and it will be true. Meta's 220 percent is true.
The question worth asking, of your employer or of yourself, is the second one: how much of that reached anyone, and what did it cost to keep upright. Meta had the discipline to ask it and the willingness to act on an answer it did not want. Most organisations running AI pilots right now are not tracking a 36 percent line at all, which means they cannot yet know whether they are in the same position.
For an individual using an AI assistant day to day, the practical read is milder and fairly encouraging. The things these tools are strongly good at are the things Meta's engineers were producing more of: drafting, summarising, explaining, translating, working through a problem, getting a first version onto the page. Where they remain weak is owning an outcome unsupervised, which is exactly the boundary the 36 percent figure describes. Using an assistant well means putting it on the first kind of work and keeping your own judgement on the second. Having a range of models available for different tasks helps with the first half of that; nothing helps with the second half except doing it.
Meta ran the experiment properly, at real expense, on itself. Then it published nothing, and the results reached us anyway because reporters went and got them. What Project OT produced in the end was a measurement most organisations have never taken and would not enjoy taking. Take it anyway, wherever you work.