THE TAKEAWAY
A 2× claim needs a named metric and comparable baseline. Twice the accepted briefs per hour, half the cost per brief and twice the qualified meeting rate are different claims.
The decision this guide helps you make
What would justify saying an agentic process performs twice as well?
You will leave with: A decision worksheet comparing accepted throughput, cost per accepted output, qualified meeting rate, incremental pipeline, with evidence and an accountable next step.
Start here: Name the metric and denominator.
Download this guide’s decision worksheetWrite the claim before the test
Specify population, task, baseline, intervention, metric and window. Accepted briefs per total delivery hour is one proposed metric, not an Outsell result. State acceptance and exclusions before looking at outcomes.
Throughput ratio equals agent-assisted accepted units per hour divided by baseline accepted units per hour. Cost improvement divides baseline cost per accepted unit by the agent-assisted cost. Conversion compares qualified outcomes per eligible account.
Protect comparability
Use the same brief, access, rubric and eligibility rules. Randomly allocate tasks or accounts when feasible. For matched comparisons, match starting conditions and state residual differences. Avoid contact spillover between conditions.
Record seller experience, existing relationships and account maturity. A warmer new queue cannot fairly be compared with a historical cold queue. Seasonality and market changes weaken historical causal comparisons.
Observational methods often did not recover the effects found by randomised experiments. Consumer advertising on one platform. This informs measurement design, not the effect size of an enterprise ABM programme. Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky (2019): A Comparison of Approaches to Advertising Measurement.
The practical workflow
- Name the metric and denominator
- Assign comparable work
- Freeze acceptance criteria
- Include review and correction costs
- Report uncertainty
Compare the approaches
| Approach | Useful when | Limitation | Next action |
|---|---|---|---|
| Accepted throughput | Repeatable delivery tasks | May miss commercial value | Count all hours and rejections |
| Cost per accepted output | Operating efficiency | Setup allocation changes the ratio | Disclose labour and tools |
| Qualified meeting rate | Account activation | Account mix and seller effects | Define eligible accounts |
| Incremental pipeline | Commercial effectiveness | Long windows and sparse outcomes | Use a defensible counterfactual |
Measure the complete delivery cost
Include generation, review, retries, exception handling and corrections. Count rejected output. Disclose labour, tools and allocation of setup cost. Faster unusable work does not represent more accepted capacity.
Freeze criteria for correctness, source support, relevance and usability. Review accepted and rejected work, using blinded assessment where practical. Resolve reviewer disagreements before expanding the workflow.
Separate attribution from incrementality
A meeting after a message is an observed sequence, not proof that the message caused it. The advertising-measurement research illustrates why observational estimates can disagree with experiments. Its setting is consumer advertising; its application here concerns test design.
Choose an enterprise observation window that fits the buying cycle. Report independent account counts, outcomes and uncertainty. If the sample cannot detect a commercial difference, describe the operational finding and continue observation.
Within the tested AI capability boundary, participants completed 12.2% more tasks and finished 25.1% more quickly. Outside it, correctness fell by 19 percentage points. Consulting tasks with GPT-4 in 2023; final paper published March 2026. Task productivity is not pipeline growth. Fabrizio Dell’Acqua and co-authors (2026): Navigating the Jagged Technological Frontier.
Publish an inspectable result
Include the dated protocol, baseline, quality threshold, total cost and limitations. Obtain permission for client identifiers and outcomes. Preserve definitions behind the summary.
The research-library calculator explores user-entered inputs. Its arithmetic is deterministic and its example illustrative. A calculated ratio is not an experimental estimate or a verified agency comparison.
Your next-action checklist
- Accepted throughput: Count all hours and rejections. Check the limitation: may miss commercial value.
- Cost per accepted output: Disclose labour and tools. Check the limitation: setup allocation changes the ratio.
- Qualified meeting rate: Define eligible accounts. Check the limitation: account mix and seller effects.
- Incremental pipeline: Use a defensible counterfactual. Check the limitation: long windows and sparse outcomes.
Use the comparison to choose a bounded next step. Record the evidence, the responsible owner, and the review decision before extending the play to additional accounts.
How to use the evidence
Read each reference against the claim it supports. Platform documentation describes capabilities; public cases report a publisher’s experience; research findings apply to the studied task and population. The workflow in this guide is an operating proposal to evaluate in your own account context.
Inspect the research library and connect this guide to measurement and revenue operations.
Questions this guide answers
What would justify saying an agentic process performs twice as well?
A 2× claim needs a named metric and comparable baseline. Twice the accepted briefs per hour, half the cost per brief and twice the qualified meeting rate are different claims.
What should I do first?
Name the metric and denominator. Record the input evidence and the acceptance criteria before continuing. Use the decision worksheet to document the owner, review date and next action.
Sources and further reading
The links below support the specific technical or platform points described here. The operating frameworks and scenarios are illustrative guidance.
- Brett Gordon, Florian Zettelmeyer, Neha Bhargava and Dan Chapsky (2019): A Comparison of Approaches to Advertising MeasurementObservational methods often did not recover the effects found by randomised experiments. Consumer advertising on one platform. This informs measurement design, not the effect size of an enterprise ABM programme.
- Fabrizio Dell’Acqua and co-authors (2026): Navigating the Jagged Technological FrontierWithin the tested AI capability boundary, participants completed 12.2% more tasks and finished 25.1% more quickly. Outside it, correctness fell by 19 percentage points. Consulting tasks with GPT-4 in 2023; final paper published March 2026. Task productivity is not pipeline growth.
Connect this guide to the next decision
Evaluate Account Cohorts at Comparable Maturity — How can an enterprise ABM team compare account cohorts without letting selection and follow-up differences dictate the result?
Run ABM Experiments With Few Independent Accounts — What can a small account-based experiment establish, and how should teams design it before seeing results?
What research actually proves about AI in marketing work — Which AI findings can inform an enterprise ABM pilot?
PUT IT INTO PRACTICE
Start with your account priorities.
Compare account focus, personalisation, deliverables, and measurement.
Explore Momentum