THE TAKEAWAY
Compare complete task designs. Human involvement helps when it supplies useful judgment; an undefined review step can add time without correcting the right errors.
The decision this guide helps you make
When should an ABM task be handled by a person, an agent or both?
You will leave with: A task-allocation matrix with evidence-based review responsibilities.
Start here: Decompose the task.
Download this guide’s decision worksheetSeparate accountability from repetitive work
Accountability means owning the decision and its consequences. It does not mean a specialist must manually complete every repeatable step. An agent can retrieve documents, reconcile a structured field or prepare a draft while a named owner remains responsible for acceptance and activation.
Break the task into observable parts before choosing the team structure. Entity matching, claim verification, stakeholder interpretation and customer contact need different inputs and error checks. A single label such as human-led or agent-first does not describe how those decisions are made.
What the meta-analysis found
Vaccaro, Almaatouq and Malone reviewed 106 experiments containing 370 effect sizes. Their 2024 meta-analysis found that combined human–AI performance was lower than the better standalone participant on average, with substantial variation by task. Creation tasks showed different patterns from decision tasks.
This does not establish that all review is harmful or that all autonomous work is superior. The evidence combines different systems, tasks and collaboration designs. Our application is to test the actual combination used for an account decision instead of assuming that adding a reviewer automatically improves it.
Explore the original methods and findings in When combinations of humans and AI are useful.
The practical workflow
- Decompose the task
- Compare three work designs
- Assign a specific review decision
- Measure corrections and total time
- Revisit allocation on exceptions
Compare the approaches
| Approach | Useful when | Limitation | Next action |
|---|---|---|---|
| Human-only | Specialised contextual judgment | Capacity may be limited | Document the decision inputs |
| Agent-only test | Repeatable bounded work | May fail on exceptions | Evaluate in a controlled setting |
| Combined design | Complementary evidence and judgment | Review can add friction | Give reviewers a precise task |
| Deterministic check | A rule can be specified | Cannot interpret every exception | Escalate unresolved cases |
Give review a specific job
A claim reviewer checks whether a passage supports an assertion. An account owner tests whether the hypothesis fits a known relationship. An activation approver checks whether contact is appropriate and permitted. Assigning one person to casually inspect the entire draft can leave these responsibilities unclear.
Show reviewers the source and the precise decision they need to make. Preserve disagreements rather than forcing approval through an attractive summary. Make it possible to return a draft for a named correction. Review should create a better accepted output, not simply a record that someone clicked approve.
Compare three versions of one brief
For a fictional account task, prepare human-only, agent-only and combined outputs using the same source collection and acceptance rubric. Include preparation, review and correction time. Have a separate evaluator assess identity accuracy, support and usefulness without seeing the process label where practical.
The agent-only output can be evaluated in a sandbox without being sent to buyers. Operational accountability remains in place during the test. If the combined design catches mistakes but doubles correction time, inspect the interface and division of work before dismissing either participant.
Look for complementary strengths
Noy and Zhang’s professional-writing experiment provides a useful example of measured assistance, but its outcomes concern writing tasks rather than ABM buying decisions. Keep that distinction when choosing which steps to automate.
Use repeated structured checks for conditions that software can enforce consistently. Use specialised account judgment where the available evidence is incomplete and relationship context matters. Track which reviewer interventions change acceptance. If a review never changes a material decision, investigate whether it is redundant or whether the reviewer lacks the information needed.
Explore the original methods and findings in Experimental evidence on the productivity effects of generative artificial intelligence.
Revisit allocation as the task changes
A routine account-data update can become an exception when two sources disagree. A low-risk draft can become consequential when it includes a sensitive claim. Route these changes explicitly rather than using the same review depth for every output.
Record the task, acceptance owner, required evidence, automatic checks and escalation condition. Evaluate that matrix after material changes in model capability, source access or programme scope. The operating objective is reliable accepted work at a defensible total cost.
Your next-action checklist
- Human-only: Document the decision inputs. Check the limitation: capacity may be limited.
- Agent-only test: Evaluate in a controlled setting. Check the limitation: may fail on exceptions.
- Combined design: Give reviewers a precise task. Check the limitation: review can add friction.
- Deterministic check: Escalate unresolved cases. Check the limitation: cannot interpret every exception.
Use the comparison to choose a bounded next step. Record the evidence, the responsible owner, and the review decision before extending the play to additional accounts.
How to use the evidence
Read each reference against the claim it supports. Platform documentation describes capabilities; public cases report a publisher’s experience; research findings apply to the studied task and population. The workflow in this guide is an operating proposal to evaluate in your own account context.
Inspect the research library and connect this guide to agentic operations.
Questions this guide answers
When should an ABM task be handled by a person, an agent or both?
Compare complete task designs. Human involvement helps when it supplies useful judgment; an undefined review step can add time without correcting the right errors.
What should I do first?
Decompose the task. Record the input evidence and the acceptance criteria before continuing. Use the decision worksheet to document the owner, review date and next action.
Read the original research
The guide explains the findings above. Open a publication to inspect its methods, setting and qualifications.
When combinations of humans and AI are useful. The pooled result covers heterogeneous tasks and earlier systems, not one ABM workflow.
Experimental evidence on the productivity effects of generative artificial intelligence. Professional writing tasks are different from factual account research and qualified meetings.
Connect this guide to the next decision
Make Human Approval a Specific Business Decision — Where should human approval enter an ABM agent workflow, and what must the reviewer see to make it meaningful?
AI writing for ABM: distinguish better drafts from better meetings — How should a revenue team evaluate AI-generated messages?
How to test a 2× ABM performance claim — What would justify saying an agentic process performs twice as well?
PUT IT INTO PRACTICE
Start with your account priorities.
Compare account focus, personalisation, deliverables, and measurement.
Explore Momentum