The talk of productivity has been something of a sideshow. What is actually happening is a change in how product management is done: it is now about providing context for a collaborator that does not sleep and has no memory of its own.

Let's recall the meeting that was once the calling card of the PM’s week. The manager would come in with a document put together over three days, only for the group to spend forty minutes realising it was already obsolete. Today the document is done in ninety seconds. The meeting goes on as before. In some ways nothing of consequence has changed.

That is the hard truth behind all the current excitement over AI in our field. We have put the artefact on autopilot but left the work as it was. The PRD composes itself, user stories are generated in droves, and Jira gets clogged up faster than one can groom it. These are genuine time savers, certainly, but they are not transformative; they are the software world’s version of a swifter printer.

One need not look far for evidence. McKinsey looked at close to 2,000 organisations in 105 countries and saw that while 88% have AI in some function, only 39% can point to an EBIT impact across the enterprise. Those few who are getting value from it do so by being willing to overhaul their workflows, not just put a new face on them. Google’s DORA report from 2025, which polled almost 5,000 tech professionals, is even more direct: AI will not repair a broken team, it will only magnify what is there. If the team is in disarray, it will be in disarray more quickly, albeit with neater formatting. A good team simply becomes more precise.

So the kind of transformation product managers have been expecting is not going to come from a superior document generator. It comes from having a collaborator on hand for the whole lifecycle, from discovery and competitive analysis to roadmap debates and post-launch monitoring. The role of the PM then becomes one of curating the judgement and context that give any artefact its worth, rather than churning out the artefacts themselves.


A tactical trap

There is nothing to it when one sits down to write a PRD. The difficulty lies in the four weeks of uncertainty that precede it: the disputes over what segment is of any consequence, an analytics event left uninstrumented, or the engineer who puts paid to the notion that the plan is workable. In the end the document is no more than a receipt for thinking already done. You cannot put the thinking on autopilot by automating the receipt.

It can also be deceiving, giving the impression of forward motion where there is none. Take the randomised controlled trial METR put before some veteran open-source developers in their own repositories. They were confident AI would put 24% to their advantage and, after the fact, still thought they had gained 20%. In reality, they were 19% slower. METR has since revised the study design and views these numbers as a snapshot of the tooling available in early 2025, not a final word. Yet the lesson remains: what looks like velocity is not always the thing, and one should have little faith in output volume as a measure of progress.

For a PM this ought to be unsettling. A developer’s work is easier to put a number on; a feature ships with a merge commit, whereas a sound call on prioritisation leaves no such mark. If a developer can misread his throughput on something with hard criteria, a PM will do no better when there are none. There is a temptation to see leverage in churning out four times the user stories, but it is often the reverse, only adding to the review burden and surface area. DORA would point to the failure mode: any speed an individual picks up is easily lost to the disorder of testing and deployment further down the line, creating pockets of productivity the rest of the system does not benefit from.


What a teammate does, tool does not

There is an architectural, not a sentimental, difference between what a tool will do and what a teammate does. You call on a tool for an output and it forgets the rest; a teammate has a memory of context and will be part of phases you have not explicitly asked it to be in.

One could make a case that today’s systems are encroaching on that territory. Researchers from Harvard and Wharton put this to the test in a field experiment at Procter & Gamble with 776 professionals facing off against actual innovation problems. They saw people using AI perform on par with two-person teams operating without it. But the real story was in how the work was done. R&D types would put forward technical answers and commercial specialists their own, as is the way without AI. Put the AI in the room, and both came up with integrated proposals that mixed the two viewpoints, erasing the functional silo. The researchers have a good name for it: the cybernetic teammate. Since bridging the gap between technical and commercial is much of what a PM is there for, a system that makes a habit of it does not eliminate the PM so much as his or her monopoly on the translation.

Yet the line is not always clear. A study by Harvard and BCG of 758 consultants showed as much. When the task was well within the model’s wheelhouse, those with GPT-4 were 25.1% quicker and turned in better work, finishing 12.2% more of it. Step outside that competence and the same consultants were 19% points behind their non-AI colleagues in getting to the right answer. They refer to it as the jagged technological frontier, and the term is apt. Two tasks can appear of equal difficulty but fall on either side of that divide with no indication from the model as to which is which. In an environment where AI is putting pen to paper on everything, understanding where that frontier lies in your domain is hardly an ancillary skill. It is practically the job.


For the product manager, context has become the artefact of choice

One might ask what a PM is putting out there if document production is no longer the main task. The answer is to be found in the practices developed to get agents to perform; at Anthropic, for instance, the engineers have put a name to it. They call it context engineering and see it as the logical evolution of prompt work. It is a matter of curating the right information for the model at inference time, with an understanding that context is a finite commodity and not something to be had for free. There is a technicality to the way they frame it, but the onus it places on the PM is anything but.

Take a roadmap conversation with an AI partner. For the tool to have any value, it must know why the company walked away from a particular segment eighteen months back, or which architectural limitations put a price on an otherwise simple feature. It should be aware of the two experiments that put this hypothesis to the test and came up short. You will not find any of that in the warehouse or the tracker or the code. It is held by four people, and two of them are gone.

Maintaining that context is emerging as the highest-leverage thing a PM does. It is not documentation in the old sense, because its primary consumer is not a human onboarding next quarter. It is the substrate determining whether every AI contribution is grounded or invented. A team with well-captured context gets a collaborator with institutional memory. A team whose context lives in Slack threads gets a fluent intern who confidently repeats mistakes the company made in 2023.

This reframes unglamorous work as strategic. The decision log stops being bureaucracy and becomes ground truth for every future conversation. The statement of what the product deliberately does not do becomes a live constraint the AI can respect. DORA's finding that platform quality determines whether AI adoption helps or does nothing has a product-side analogue: context quality determines whether an AI collaborator compounds your judgement or dilutes it.


Across the lifecycle, not at the end of it

The value is in the ongoing engagement over the full lifecycle, not at its conclusion. You will find that pattern to be true in every phase: what matters is a steady hand at the wheel, not a single act of generation.

Consider discovery and research. Here, the ability to generate is of little note; any model will do with summarising thirty interview transcripts. The work is transformed by a collaborator who can have all thirty at hand as well as two years’ worth of churn interviews and support tickets, and put them to you in real time. “We have already heard and put aside that complaint,” it might say, or “what would need to be the case for your conclusion to be wrong?” That is no mere summary, but a research partner challenging your reading of the data.

Then there is competitive analysis, where the task aligns with new agent architectures in an unusual way. In one instance, Anthropic detailed a multi-agent set up for its research system: a lead agent lays out a plan and sets off subagents to look into separate angles, with each coming back with their findings. On the books, that approach beat a like-for-like single agent by 90.2% on breadth-first tasks. They are open about the price of that performance, though; token usage is some fifteen times what a chat would take, and they caution against using it where agents are too interdependent or need to share context. But when it comes to mapping out the competition, the job breaks down into independent threads quite nicely. Figuring out what your product ought to do in light of that does not.

In experimentation, AI's contribution is best understood against a sobering base rate. Only about a third of ideas tested on Microsoft's platform enhanced metrics they were designed to improve, and in optimised domains the rate is worse. The implication is not that AI generates more hypotheses - idea supply was never the bottleneck. It is that a persistent collaborator holds the history of what has been tried, flags when a proposed test is a 2024 experiment renamed, and pressure-tests whether the success metric captures what the team cares about. Two-thirds of experiments failing is the cost of learning; running the same failed experiment twice is a context problem.

In technical trade-off evaluation, the change is that PMs can inspect reality directly. Anthropic's guidance on adopting agentic coding notes that product managers can now write requirements informed by actual codebase constraints, and its internal case study describes teams treating the coding agent as a first stop for identifying which files a change would touch, eliminating the manual context-gathering that used to precede any estimate. A PM who can ask what a change would require and get a grounded answer before sprint planning is not doing engineering. They are asking better questions, which is the job.

In release and post-launch monitoring, persistence is the whole value. The characteristic failure of product management is not launching the wrong thing; it is launching the right thing and losing interest during the eight weeks when evidence accumulates across five systems. A collaborator that watches funnel metrics, support volume, and reviews against the hypothesis the feature was built to test and raises a hand when they diverge closes a loop human attention has always closed badly.


A look at the multi-agent product team and where it falls short

These days one can put in place a product function with agents taking on separate hats: engineer, designer, QA, analyst, knowledge manager or researcher. There is no speculation to it. You will find that division of labour in the way Anthropic’s own teams operate, with design staff turning to agents for pull-request reviews and test coverage, non-technical people formulating data queries in plain English and engineers putting through several persistent instances of an agent to keep context over days of work in parallel. The numbers from Stanford HAI back up the capability curve; they have seen agent performance on OSWorld, a measure of real computer tasks, go from 12 to some 66% in the course of a year.

Yet there is a counterargument in those very figures. A two-thirds hit rate leaves one in three tries to fail, and not in any random fashion between the hard ones and the easy ones. The data also puts a fine point on the oddities at the frontier: a model might take home gold in olympiad maths, yet only half the time will it read an analogue clock right.

Orchestration, then, is a matter for humans. It takes a person to determine what can be broken down into parallel threads and what is too interdependent or to step in when two agents are at odds. Someone has to be the one to see that six coherent outputs from as many agents do not add up to a product anyone is interested in. Most have yet to develop that aptitude. McKinsey reports 62% of firms are dabbling with agents, but just 23% have managed to scale them in any single function. For the ones that are making headway, it comes down to how they have reorganised, not simply having access to the models.


The judgement problem

It is worth putting one risk on the table in plain terms, since it is a quiet thing that runs counter to any sense of forward momentum. A study by Microsoft Research and Carnegie looked at 319 knowledge workers in 936 actual cases of AI deployment and turned up an unambiguous inverse correlation: as confidence in the AI went up, the critical scrutiny of its output went down. Put more trust in your own skills and the opposite was true. The subjects said their mode of thinking had changed too, with less time spent gathering information and more put into verification and seeing a task through.

The authors of the paper are drawing on something well known in automation circles: when you mechanise the routine and reserve only the exceptions for people, you rob them of the practice necessary to maintain good judgement. Consider a product manager who has not had to endure a hard-fought prioritisation debate; he may be none the wiser if an AI has put together a neat ranking that is optimised for the wrong outcome. Or the PM who has never penned a poor PRD will lack the instinct to spot a subtle flaw in one. You have to make your own mistakes to develop the scepticism required to identify a wrong answer that is otherwise well put together.

In practical terms this goes against the grain of efficiency. There is some friction that is essential. To form an opinion of your own before turning to the model on the decisions of consequence is no act of stubbornness but a way of staying calibrated to what is returned. Good teams know how to handle this; they determine ahead of time which matters will be given the fast track and which require the slower approach.


What this actually looks like on Monday

The change is not so much in the tools one puts in place as in what you consider to be your output. There are four areas where that difference is most pronounced.

To start, take ownership of product context and view it as an asset to be tended to. You need a structured place for your decisions, the constraints you face, dead ends from experiments and the rationale for the strategy at hand, not to mention any alternatives put aside. Make sure all of that is retrievable since an AI is only as useful as the information it has access to. If nothing else is done, do this.

Then there is the matter of your own jagged frontier. Keep a map of it. Put in some time every quarter to run your regular work — say a feasibility scoping or some competitive analysis — both with and without AI assistance. Note down where the machine was dependable and where it was wrong, especially the errors it made with such confidence they were easy to overlook. The frontier shifts with every new model, and a map from half a year ago is a liability in as much as it gives you a false sense of knowledge.

You will also find your review burden shifting from the finished piece to the framing of it. Let the AI put together the draft; your worth is in the question that elicits it. In an AI workflow, the first minute is the one with the most leverage, so put in the effort on the problem statement rather than polishing the prose after the fact.

Finally, get ahead of the issue of accountability. Before an agent has a chance to triage feedback or flag a regression, determine who is on the hook for the result and what line will compel a human to step in. There is no technical fix for the uncomfortable reality of who answers for an agentic recommendation that causes harm; it is an organisational matter. Those teams that have not come to terms with it will be forced to under less-than-ideal circumstances.


What is left to the human

It is tempting to end on the reassuring note that AI handles mechanical work while humans do visionary work. It is a flattering way to put it, but not entirely true. The more difficult aspects of product management have never been purely mechanical, and there is no mistaking the fact that AI can now do some of the things we once put in the human column.

The distinction is more fundamental and narrower than that. An AI system will produce an output, but it has no skin in the game. It can lay out the trade-off of shipping in March with three defects versus waiting until June for a clean release. But it cannot make the call on which future the company ought to inhabit, since a decision requires being answerable for the consequences that will be felt by actual people. It can report back on what a thousand customers have said; it cannot tell you which of them you have made up your mind to build for. That is not a matter of fact so much as a statement of who you intend to be.

The evidence is clear. You find value in workflows that have been rethought from the ground up, not just put on a fast track. The machine’s capabilities are uneven, something only a domain expert can work around. Put too much faith in it and you lose the kind of scrutiny that makes it safe to use. In the end, context is the limiting factor: a working knowledge of the product, its audience, what has been done before and what the organisation has chosen to prioritise.

That is where the job lies these days. Not in putting pen to paper. It is in holding the context, framing the issue, posing the question that changes the tenor of the debate, and making sure the AI does not stray from what the business is set to accomplish. Then you put your name to the decision. The AI is a capable collaborator and one that is getting better at a pace. But it has no concept of what is important. That will always be yours to provide.