THE EVIDENCE, IN CONTEXT

Prove the work.
Then scale it.

Research for enterprise teams evaluating AI-assisted account work. See what was measured, where the finding applies, and what still needs a test.

Curated by Outsell AI · Evidence reviewed · Source policy

RESEARCH EXPLAINED

Understand the finding. Apply it carefully.

Read practical explanations of retrieval, personalisation, B2B account work and causal measurement. Each guide connects the original research to a worked example and a decision worksheet.

Choose a learning path

Inspect the original evidence.

Open a study’s context for its setting, limits and a proposed ABM application. Research, surveys and technical guidance are labelled separately.

NeurIPS research paper · 2020

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis and co-authors

Retrieval-augmented models produced more factual generation than the tested parametric-only baseline.

Study context and ABM application
Study setting
Knowledge-intensive question answering and generation benchmarks
Boundary of the evidence
Wikipedia-based benchmark results do not establish accuracy for a private account database.
What we would test in ABM
Test retrieval and claim support separately before accepting a brief.

EMNLP research paper · 2023

Enabling Large Language Models to Generate Text with Citations

Tianyu Gao, Howard Yen, Jiatong Yu and Danqi Chen

The study evaluates fluency, correctness and citation quality as separate dimensions.

Study context and ABM application
Study setting
ALCE question-answering and citation benchmark
Boundary of the evidence
A plausible citation can still fail to support the sentence beside it.
What we would test in ABM
Check whether each material assertion is entailed by its cited passage.

Peer-reviewed TACL research paper · 2024

Lost in the Middle: How Language Models Use Long Contexts

Nelson F. Liu and co-authors

Performance depended on where relevant information appeared in the tested contexts.

Study context and ABM application
Study setting
Multi-document question answering and key-value retrieval
Boundary of the evidence
The tested model generation does not determine the behaviour of every current model.
What we would test in ABM
Move critical constraints around a representative brief and test recall.

ICLR research paper · 2023

ReAct: Synergizing Reasoning and Acting in Language Models

Shunyu Yao and co-authors

Interleaving reasoning and actions improved results on the studied benchmarks.

Study context and ABM application
Study setting
Question answering, fact verification and interactive task benchmarks
Boundary of the evidence
Benchmark success is not permission to operate a production marketing system autonomously.
What we would test in ABM
Design observable tool steps with bounded actions and stop conditions.

Security research paper · 2023

Not what you have signed up for: Indirect Prompt Injection

Kai Greshake and co-authors

Instructions embedded in retrieved data could redirect application behaviour.

Study context and ABM application
Study setting
Demonstrations against LLM-integrated applications
Boundary of the evidence
The paper demonstrates attack mechanisms rather than a current marketing incident rate.
What we would test in ABM
Treat retrieved pages as data and enforce tool permissions outside the model.

Peer-reviewed systematic review and meta-analysis · 2024

When combinations of humans and AI are useful

Michelle Vaccaro, Abdullah Almaatouq and Thomas W. Malone

Human–AI combinations underperformed the better standalone participant on average; outcomes varied by task.

Study context and ABM application
Study setting
106 experiments; 370 effect sizes
Boundary of the evidence
The pooled result covers heterogeneous tasks and earlier systems, not one ABM workflow.
What we would test in ABM
Compare human, agent and combined versions of the same task.

Peer-reviewed randomised writing experiment · 2023

Experimental evidence on the productivity effects of generative artificial intelligence

Shakked Noy and Whitney Zhang

Average task time fell 40% and assessed writing quality rose 18% with ChatGPT.

Study context and ABM application
Study setting
453 college-educated professionals
Boundary of the evidence
Professional writing tasks are different from factual account research and qualified meetings.
What we would test in ABM
Measure accepted message quality and total production time before testing response.

Peer-reviewed marketing research · 2015

Unraveling the personalization paradox

Elizabeth Aguirre and co-authors

Responses to personalised advertising depended on how information was collected and on trust cues.

Study context and ABM application
Study setting
Field evidence and three consumer advertising experiments
Boundary of the evidence
Consumer advertising findings do not establish enterprise reply rates or legal compliance.
What we would test in ABM
Use relevant business context without suggesting undisclosed personal surveillance.

Peer-reviewed advertising field experiment · 2013

When Does Retargeting Work? Information Specificity in Online Advertising

Anja Lambrecht and Catherine Tucker

Specific retargeted ads were less effective on average; preference development changed their usefulness.

Study context and ABM application
Study setting
Online travel advertising field experiment
Boundary of the evidence
Travel-product advertising is not an enterprise software procurement experiment.
What we would test in ABM
Match message specificity to observed evaluation readiness.

Peer-reviewed B2B survey research · 2025

The influence of key account management on competitive advantage and firm performance

Farbod Fakhreddin, Pantea Foroudi and Kaouther Kooli

Survey modelling connected key-account orientation, capabilities, competitive advantage and performance.

Study context and ABM application
Study setting
568 European B2B supplier firms
Boundary of the evidence
Associations in a survey cannot establish the causal return of an ABM campaign.
What we would test in ABM
Inspect relationship capability and delivery ownership alongside campaign execution.

Publisher-reported B2B buyer survey · 2026

The surprising economics of B2B growth

McKinsey & Company

Buyers reported using an average of ten channels across the purchasing journey.

Study context and ABM application
Study setting
Nearly 4,000 decision-makers across 13 countries
Boundary of the evidence
Reported survey patterns are not a causal estimate of adding channels to your programme.
What we would test in ABM
Carry the same decision context across the channels your buyers actually use.

Peer-reviewed advertising measurement research · 2017

Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness

Garrett A. Johnson, Randall A. Lewis and Elmar I. Nubbemeyer

Ghost ads identify control counterparts of exposed users to improve experimental measurement.

Study context and ABM application
Study setting
Randomised advertising experiments with counterfactual exposure records
Boundary of the evidence
This requires platform support; ordinary CRM records cannot recreate ghost-ad allocation.
What we would test in ABM
Preserve assignment and eligibility before interpreting campaign exposure.

ICML research paper · 2017

On Calibration of Modern Neural Networks

Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Q. Weinberger

Classification confidence and actual correctness were misaligned; post-processing could improve calibration.

Study context and ABM application
Study setting
Image and document classification experiments
Boundary of the evidence
These datasets do not validate a particular ABM propensity model.
What we would test in ABM
Test whether score bands correspond to observed outcomes in later cohorts.

Peer-reviewed evaluation-method research · 2015

The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets

Takaya Saito and Marc Rehmsmeier

Precision–recall exposes the reliability of positive predictions under class imbalance.

Study context and ABM application
Study setting
Imbalanced binary-classification examples
Boundary of the evidence
Precision depends on prevalence; comparisons need the same evaluation population.
What we would test in ABM
Report useful accounts found within the sales team’s available capacity.

Technical specification overview · 2013

PROV-Overview

World Wide Web Consortium

PROV represents relationships among entities, activities and responsible agents.

Study context and ABM application
Study setting
The W3C PROV family of provenance specifications
Boundary of the evidence
This is a data specification, not an experiment demonstrating commercial performance.
What we would test in ABM
Record the source, transformation and owner behind an account assertion.

Peer-reviewed statistical methods paper · 2019

Metalearners for estimating heterogeneous treatment effects using machine learning

Sören R. Künzel, Jasjeet S. Sekhon, Peter J. Bickel and Bin Yu

The paper develops approaches for estimating how treatment effects differ across contexts.

Study context and ABM application
Study setting
Methods and examples for conditional treatment-effect estimation
Boundary of the evidence
Causal identification assumptions and adequate data remain necessary.
What we would test in ABM
Distinguish likely converters from accounts whose outcomes contact might change.

Cross-sectoral risk-management guidance · 2024

Generative Artificial Intelligence Profile: NIST AI 600-1

National Institute of Standards and Technology

The profile describes generative-AI risks and proposed management actions.

Study context and ABM application
Study setting
Companion profile to the AI Risk Management Framework
Boundary of the evidence
Guidance is not a product certification or an ABM performance benchmark.
What we would test in ABM
Connect material failure modes to owners, checks and monitored evidence.

Peer-reviewed marketing research synthesis · 2016

Understanding Customer Experience Throughout the Customer Journey

Katherine N. Lemon and Peter C. Verhoef

The review connects experience over time with multiple touchpoints and organisational functions.

Study context and ABM application
Study setting
Synthesis of customer-experience and journey research
Boundary of the evidence
A research synthesis does not provide a fixed enterprise conversion formula.
What we would test in ABM
Design evidence and ownership around a buyer task across touchpoints.

Peer-reviewed field study · 2025

Generative AI at Work

Erik Brynjolfsson, Danielle Li and Lindsey Raymond

AI assistance increased issues resolved per hour by 15% on average.

Study context and ABM application
Study setting
5,172 customer-support agents
Boundary of the evidence
Customer support with substantial differences between workers. This is not an ABM revenue lift or an Outsell result.
What we would test in ABM
Test accepted account work per hour, including verification and rework.

Peer-reviewed randomised experiment · 2026

Navigating the Jagged Technological Frontier

Fabrizio Dell’Acqua and co-authors

Within the tested AI capability boundary, participants completed 12.2% more tasks and finished 25.1% more quickly. Outside it, correctness fell by 19 percentage points.

Study context and ABM application
Study setting
758 BCG consultants
Boundary of the evidence
Consulting tasks with GPT-4 in 2023; final paper published March 2026. Task productivity is not pipeline growth.
What we would test in ABM
Evaluate extraction, drafting and strategic inference separately.

Randomised controlled trial · 2025

Early-2025 AI and Experienced Developer Productivity

Joel Becker, Nate Rush, Beth Barnes and David Rein

Developers took 19% longer with the early-2025 AI tools in the studied repositories.

Study context and ABM application
Study setting
16 experienced developers; 246 tasks
Boundary of the evidence
A narrow coding setting. A February 2026 follow-up reports possible speedups but serious selection and measurement limits. This is not a timeless verdict on AI.
What we would test in ABM
Measure elapsed work rather than whether automation feels faster.

Peer-reviewed comparison with field experiments · 2019

A Comparison of Approaches to Advertising Measurement

Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky

Observational methods often did not recover the effects found by randomised experiments.

Study context and ABM application
Study setting
15 US Facebook advertising experiments
Boundary of the evidence
Consumer advertising on one platform. This informs measurement design, not the effect size of an enterprise ABM programme.
What we would test in ABM
Separate touchpoint attribution from incrementality; use account-level holdouts where feasible.

KDD research paper and benchmark · 2024

GEO: Generative Engine Optimization

Pranjal Aggarwal and co-authors

Content interventions improved the benchmark’s visibility measure by up to 40%, with effects varying by domain.

Study context and ABM application
Study setting
GEO-bench across multiple query domains
Boundary of the evidence
A benchmark result, not a forecast of traffic, live AI citations or Google rankings for this website.
What we would test in ABM
Make answers extractable and verifiable, then observe actual citations and referrals.

HICSS conference paper; critical discourse analysis · 2022

AI Agents, Humans and Untangling the Marketing of Artificial Intelligence in Learning Environments

Isabel Pedersen and Ann Hill Duin

The paper examines how promotional language presents AI agency and human roles in learning.

Study context and ABM application
Study setting
Corporate representations of AI in learning environments
Boundary of the evidence
It is not an agency performance experiment and provides no Outsell-versus-competitor or 2× result.
What we would test in ABM
Describe the actual agent task, evidence and decision owner.

YOUR INPUTS. EXPLICIT MATH.

Compare accepted work, not generated volume.

Use comparable tasks and the same quality threshold. Include review, retries and correction in total hours and cost. The prefilled values are an illustrative example, not an Outsell result.

Baseline process
Agent-assisted process
How the calculation works

Throughput ratio = (assisted accepted outputs ÷ assisted hours) ÷ (baseline accepted outputs ÷ baseline hours). Cost efficiency = baseline cost per accepted output ÷ assisted cost per accepted output. A 2× throughput ratio and a 2× cost-efficiency ratio describe different outcomes.

This calculator does not estimate meeting conversion, revenue lift or statistical significance. Read the comparison protocol.

Turn the evidence into a decision.