AI has spread quickly across the enterprise. Measurable business value hasn’t.
While AI adoption has become widespread across most organizations, most have not yet demonstrated meaningful business impact that the CFO, CEO, or the board care about. McKinsey reports that AI adoption remains high at nearly nine in ten organizations, yet only 37% report positive enterprise-level EBIT impact resulting from that adoption. Deloitte’s research echoes this when looking specifically at agentic AI—only 10% of organizations already adopting it say they are seeing scaled, measurable ROI.
It’s no wonder that OpenAI and Anthropic have invested billions to stand up services organizations and embed forward-deployed engineers to ensure their enterprise customers get to value soon.
Many organizations have responded with a standard playbook: license a copilot for the team, encourage AI adoption, spin up a handful of pilots, appoint an AI council, and wait for the productivity to show up in the numbers. Mostly it hasn’t.
That pattern isn’t a technology failure—it’s a conditions failure. For the loop to function at scale—agents sensing, reasoning, acting, and learning on real work, with humans directing and validating—both the agents and the people need things most organizations haven’t yet built: coherent context, connected tools, content machines can use, clear decision rights, and measurement fast enough to learn from. Deploy tools without those conditions, and you get exactly what the studies describe: high adoption and low transformation.
The companies getting value are redesigning the work.
McKinsey’s research gives us a pretty clear direction. In a 2026 study of AI transformation readiness, leaders were 5.3× more likely to report enterprise value capture when workflows were redesigned than when they remained unchanged — 32% versus 6%. That said, IBM found that more than three-quarters of executives say most of their AI investment has gone toward improving existing processes. In contrast, 78% say capturing the maximum benefit from agentic AI requires a new operating model.
We know the work needs to change. But we keep investing in ways to make the old work faster.
The technology still isn’t the hard part.
I spent years leading the marketing technology stack at a large enterprise, and one truth kept repeating itself: the technology is rarely the hard part. We could buy or build a genuinely useful business capability, get it through IT and security, integrate it into the stack, and still struggle to capture the value. The hard work was the same tough slog of digital transformation: executive alignment, cross-functional ownership, operating-model changes, team buy-in, process redesign, incentives, governance, and actually getting people to change how they work. Technology can enable a different way of working. It can’t make the organization operate differently.
That hasn’t changed with AI. If anything, AI is making the old transformation problems more visible. Agents raise the stakes because they don’t just sit inside a tool waiting for someone to use them. They can research, decide, create, update systems, trigger actions, and hand work to other agents. Put that capability into a workflow with unclear ownership, contradicting approvals, bad handoffs, or missing context, and the agent inherits all of it — then starts executing it at scale.
A method for engineering human-agent work
The basic discipline isn’t new—it’s been around for a long time. Business process reengineering showed us to start with the outcome, understand the work end to end, challenge the current process, and design a better one. Agents add several elements the old playbooks didn’t have to answer in as much detail: which judgments should remain human, what context each actor needs at the moment of work, how much autonomy each part of the workflow gets, and what oversight needs to surround it?
An AI-first redesign has to work across four connected dimensions.
| Dimension | Core consideration |
| People + Agents | Who — or what — should execute the work? |
| Process + Workflow | What are the new steps and gates? |
| Data + Context | What context is required to do the work? |
| Technology + Tools | What tooling delivers automation and scale? |
Governance and risk span all four, defining what can act, what requires approval, and where human oversight belongs. Measurement points back to the question that began the whole exercise: did the redesigned workflow deliver the business outcome we started with?
We organize the approach into three phases and seven moves:
- Phase One: Prioritize the Work
- Phase Two: Engineer the Work
- Phase Three: Operationalize the Work

The quick version: break the work down, assign it, give it context, govern it, make it learn.
Prioritize the Work
1. Start from the outcome
Start one level above the AI use case. Say the GTM outcome is faster, more consistent conversion of high-intent inbound demand. One capability supporting that outcome is getting useful account intelligence into the hands of a seller quickly.
Now we have recurring work to assess: Inbound lead → account research → account brief → xDR/AE handoff.
That workflow might cross forms, first- and third-party enrichment, CRM and transaction history, website research, company research, prior engagement, custom account scoring, and some kind of seller notification. It may cross marketing and sales ownership as well. That’s enough surface area to redesign.
2. Choose the workflows
Not every process deserves an agentic workflow. Before prioritizing one, we run it through five pass/fail gates:
- Can the output be verified?
- Does the work recur often enough to matter?
- Can we put clear boundaries around what the agent can and cannot do?
- Is there a measurable outcome with someone who owns it?
- Can a human realistically review the work when review is required?
A workflow can have enormous volume, high labor cost, and good technical feasibility, but none of those make up for a use case where errors can’t be detected, or nobody is accountable for the outcome. A failed gate isn’t always rejection…it identifies the readiness work required to make the workflow eligible next time. The answer might just be…‘later.’
Assuming we make it this far, now we score what survives on two separate dimensions: value and feasibility. High-value, high-feasibility workflows become the first transformation candidates. High-value work with lower readiness may deserve a narrow experiment first, while work with low agent leverage may be a better candidate for conventional automation.
Engineer the Work
3. Break down the work
Take the workflow apart far enough to see what each step actually requires. In the inbound example, gathering company information is different from deciding whether the account matters. Summarizing engagement history is different from deciding what the seller should do. Checking an account against ICP criteria is different from deciding the initial outreach channel.
That matters when you start deciding what belongs with a person and what can move to an agent. Some decisions are rules-based and don’t require human judgement or LLM reasoning at all: does this account match the ICP, is this signal recent enough, has the required account data been captured? Others depend on direction, taste, trust, or business judgment. Those are different kinds of work, even if the current process has the same person doing all of them.
Breaking the workflow down also creates an opportunity to question why the work is structured this way in the first place. Does every step still need to exist? Can research that happens sequentially today happen in parallel? Is a handoff there because the work requires it, or because two teams happen to own different parts of the process? Does an approval address a real risk, or is it simply inherited from the way the organization has always worked?
The goal is to redesign the work before deciding where (or if) AI fits. Otherwise, it’s very easy to end up automating the same steps, handoffs, and approvals you already had.
4. Assign the work
Once the work is broken down, decide where each part belongs. We use four assignment levels:
| Assignment level | Definition |
| Human | A person performs the work. Agents may support with research, drafting, or preparation, but the human is accountable for the task and the judgment. |
| Human in the loop | The agent performs the work, but a human must review and approve it before anything consequential happens. |
| Human on the loop | The agent completes the work within defined operating boundaries. Humans monitor and intervene when predefined triggers or exceptions occur. |
| Fully agent | The agent executes autonomously within a bounded and observable domain, with metrics and reversion triggers providing oversight. |
In our inbound workflow, an agent might gather the research, build the brief, and write it to the CRM without approval, while the AE owns the interpretation and any customer outreach. The agent’s operating domain might be explicit: named B2B accounts only, approved public and internal sources, no outbound contact, no pricing claims, no changes to opportunity stage.
Those boundaries matter because autonomy is a critical design decision.
5. Engineer the context
A capable agent without the right context (or too much context) will still produce bad work. The account-brief agent may need the company and domain, ICP criteria, CRM activity, known contacts, product interest, prior opportunities, approved research sources, and the questions the seller needs answered. Write that down and treat it as a required input for that step.
The output and handoff back needs the same approach. “Strong account, recommend follow-up” might be the way the agent scores the account. But we’ve seen that without the evidence behind it, the seller may not trust it…what signals were found, where they came from, what is known, what remains uncertain, and why the agent reached its conclusion.
6. Set the guardrails
“Human in the loop” leaves a lot unanswered: where is the gate, what is the human evaluating, what happens when they reject the work, how quickly does review need to happen, and does every execution require it? The answer should follow risk and volume. High-risk actions may need blocking approval. Work inside a well-understood process may be fine for ongoing human monitoring. High-volume, lower-risk work may run independently unless a metric crosses a threshold.
There’s a new tactic here: agent-assisted oversight. An independent, adversarial review agent can monitor the work, so human attention focuses on what it escalates and on a regularly scheduled audit of what it passes. This is how a workflow can move from human-in-the-loop to human-on-the-loop more quickly. Agent-assisted oversight can extend review capacity…but it doesn’t relocate accountability: a named human owner still owns the outcome.
It should also be noted that there are really two types of governance: exercised governance, where you decide where the human gates go, and enforced governance, where the platform actually enforces them. A prompt can ask an agent to respect a given data policy…runtime controls and guardrails in the agentic platforms themselves make it impossible to cross. You need both: design defines the gates, and enforcement makes them real.
Now calculate the oversight budget. Suppose the new workflow produces 200 account briefs a month and proper review takes three minutes each. That consumes ten hours of human review. If the owner has two hours available, the workflow design is already broken.
Maybe every output gets reviewed during the pilot, and 10% get audited in production. Whatever the answer, do the math before the build, because human attention is one of the resources the workflow consumes.
Operationalize the Work
7. Turn on the loop
Run the workflow on three accounts before you run it on 200. Or even better, start with synthetic data and use an adversarial agent to review. Look at where the agent had enough context and where it guessed or hallucinated, inspect the handoffs and evidence, and pay attention to what the human had to recreate or compensate for. Then tighten the design and expand the scale and scope deliberately.
Bring the workflow back to the business outcome you started with: did high-intent leads reach sellers faster? Did conversion improve? Did seller capacity increase? Did cost fall? Did quality improve? Agent activity and business performance are different things. The redesigned workflow ultimately has to move the latter.
What compounds
A GTM organization may have dozens of important workflows, and nobody wants a two-month consulting exercise every time an agent enters one of them. Done well, you don’t start over. Each redesigned workflow leaves reusable features behind: shared context, reusable connectors, evaluation methods, approval patterns, governance controls, and a clearer model for where human judgment belongs. The next workflow should start further along than the last one.
Over time, those individual projects begin to add up to something more useful than a collection of AI implementations. You start to build a repeatable way of deciding how work should be divided across people, agents, automation, and systems…and a common set of capabilities that make the next redesign faster and easier. That’s how fifteen workflows become manageable without creating fifteen individual solutions.
The models will keep changing, the tools will keep changing, and the work agents can perform will keep expanding. The longer lasting capability is the organization’s ability to redesign the work as those capabilities evolve, while keeping business outcomes, human judgment, and accountability front and center.
That is ultimately what turns AI from a series of experiments into a new way of operating that delivers compounding business value.

What GTM Leaders are Asking
What does it mean to redesign work around AI?
It means starting with the business outcome and redesigning the workflow around the strengths of people, agents, automation, context, and systems. The goal is to decide which work should exist, who or what should perform it, what information is required at each step, where human judgment belongs, and how the workflow will be measured.
Where should a GTM organization start with agentic AI?
Start with a meaningful business outcome, then identify the recurring workflows that most directly influence it. The best first candidates are usually workflows with enough volume to matter, clear boundaries, measurable outcomes, and outputs that can be reviewed or verified. From there, prioritize based on value and feasibility.
How do you decide what work should stay human and what should move to an agent?
Break the workflow into smaller units and look at the kind of judgment each one requires. Direction, taste, trust, accountability, and high-consequence decisions often deserve human ownership. Criteria matching, research, synthesis, monitoring, and structured evaluation are often more delegable. The important distinction is the nature of the judgment, not whether the task looks difficult.
What is the difference between human in the loop and human on the loop?
Human in the loop means an agent performs the work but a person must review or approve it before anything consequential happens. Human on the loop means the agent operates within a defined boundary while people monitor performance, review samples, and intervene when specific triggers or exceptions occur.
How much autonomy should an AI agent have in a GTM workflow?
As much as the workflow can safely support. Autonomy should be designed around the consequence of errors, the ability to observe what the agent is doing, the ease of reversing mistakes, and the quality of the controls around the work. It should increase only as the workflow proves it can perform reliably.
Why does context matter so much in agentic workflows?
Agents can only make good decisions with the information available to them at the moment of work. That context may include account history, ICP criteria, product information, policies, brand guidance, customer signals, prior interactions, and approved sources. Good workflow design makes those requirements explicit rather than assuming the agent will somehow find what it needs.
What does human oversight actually look like in an agentic workflow?
It can take several forms. High-consequence work may require blocking approval before an action occurs. Higher-volume, lower-risk work may use sampling. Other work may run autonomously until a metric or exception threshold is crossed. The important part is to define the gate, the reviewer, the evidence they see, and the amount of human attention the design requires.
How do you keep human oversight from becoming a bottleneck?
Budget it before the workflow goes live. If 200 outputs require three minutes of review each, that creates ten hours of human work. If the owner only has two hours available, the design needs to change. The answer may be sampling, exception-based review, tighter operating boundaries, or moving lower-risk work to a more autonomous pattern.
How should GTM leaders measure whether a redesigned workflow is working?
Measure the business outcome, not just agent activity. Depending on the workflow, that could include pipeline conversion, cycle time, throughput, cost, seller capacity, quality, customer experience, or error rates. Agent-level measures still matter, but completing tasks successfully is not the same as creating business value.
Do we need to redesign every GTM workflow from scratch?
No. Each well-designed workflow should leave reusable assets behind: context objects, connectors, evaluation methods, governance patterns, approval logic, and operating boundaries. Over time, those shared capabilities make the next workflow easier to redesign and help the organization build a repeatable human-agent operating model.
Sources:
- McKinsey & Company, The State of AI in 2026: On the Road to ROI, August 25, 2026. Nearly nine in ten organizations report regular AI use; 37% report positive EBIT impact; eight in ten respondents report improved individual productivity. McKinsey source
- McKinsey & Company, From Adoption to Impact: Three Horizons of AI Transformation, July 8, 2026. Organizations redesigning workflows were 5.3× more likely to report enterprise value capture than those leaving workflows unchanged (32% vs. 6%). McKinsey source
- Deloitte, AI ROI: The Paradox of Rising Investment and Elusive Returns, October 22, 2025. Among organizations already using agentic AI, 10% reported significant ROI. Deloitte source
- IBM Institute for Business Value, Agentic AI’s Strategic Ascent: Shifting Operations from Incremental Gains to Net-New Impact, 2025. Research on the operating-model changes required to capture value from agentic AI. IBM source
- Accenture, Designing a New Agentic Collaborative Workforce, January 21, 2025. Accenture’s marketing redesign is expected to reduce average campaign steps from 135 to 85 and improve time to market by 25–35%. Accenture source
- Michael Hammer, Reengineering Work: Don’t Automate, Obliterate, Harvard Business Review, July–August 1990. The classic process-reengineering argument for redesigning work rather than simply automating existing processes. Harvard Business Review source

