THE TAKEAWAY

When conversions are uncommon, overall accuracy can hide missed opportunities. Report precision and recall at the team’s capacity, alongside prevalence and the cost of errors.

The decision this guide helps you make

Why can high prediction accuracy still produce a poor sales queue?

You will leave with: A capacity-based score evaluation with a transparent confusion matrix.

Start here: Define eligible accounts and outcome.

Download this guide’s decision worksheet

Start with the operating population

If only a small fraction of eligible accounts reaches a qualified outcome, a model that predicts no outcomes can look accurate. That model does nothing useful for prioritisation. Define the population and event before choosing a headline metric.

The sales decision is often which accounts to work with limited capacity. Evaluation should therefore inspect the selected queue, the useful outcomes it contains and the relevant accounts left outside it. Preserve the same population when comparing models.

What precision–recall research explains

Saito and Rehmsmeier’s 2015 PLOS ONE paper explains how precision–recall plots can make the reliability of positive predictions easier to interpret in imbalanced classification. The discussion concerns evaluation methods rather than an ABM deployment.

Our application is to show the fraction of selected accounts reaching the defined event and the fraction of all observed events captured by that selection. Precision depends on prevalence, so a result from one territory should not be transferred unchanged to another with a different account mix.

Explore the original methods and findings in The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets.

The practical workflow

Rare ABM conversions: evaluate the accounts sales can actually work. Workflow: Define eligible accounts and outcome; Count the event prevalence; Evaluate at sales capacity; Inspect false and missed positives; Validate on a later cohort.
A sequence for applying this guide. Use the review points to decide whether the work is ready to continue. View full-size image
  1. Define eligible accounts and outcome
  2. Count the event prevalence
  3. Evaluate at sales capacity
  4. Inspect false and missed positives
  5. Validate on a later cohort

Compare the approaches

Compare the approaches
ApproachUseful whenLimitationNext action
AccuracyOverall classificationCan hide rare-outcome failureShow a confusion matrix
PrecisionQuality of the selected queueChanges with prevalenceReport selected count and population
RecallCoverage of observed outcomesHigher coverage can increase workloadUse the actual capacity
CalibrationProbability reliabilityDifferent from rankingCompare predicted and observed bands
Decision guide: Rare ABM conversions: evaluate the accounts sales can actually work. Accuracy: Overall classification. NEXT ACTION: Show a confusion matrix Precision: Quality of the selected queue. NEXT ACTION: Report selected count and population Recall: Coverage of observed outcomes. NEXT ACTION: Use the actual capacity Calibration: Probability reliability. NEXT ACTION: Compare predicted and observed bands
Match the situation to a useful next action. The comparison above includes the limitations of each approach. View full-size image

Work through the arithmetic

In an illustrative cohort of 1,000 accounts, assume 20 eventually meet the outcome definition. A rule predicting no outcomes has 98% accuracy and captures none. Another rule selects 50 accounts, including 10 of the 20 outcomes. Its precision is 10 divided by 50, or 20%. Its recall is 10 divided by 20, or 50%.

These fictional values demonstrate why accuracy alone is insufficient. They do not establish the quality of an Outsell score. Report selected count, true positives, false positives and missed positives so a reader can inspect the denominators.

Choose thresholds around capacity

If the team can work fifty accounts this month, compare methods at fifty selected accounts. If research capacity differs from seller capacity, use different queues and acceptance criteria. A false positive may consume a short research task or an expensive executive workshop; those costs are different.

Do not optimise a threshold on the same cohort used to report its success. Use a separate validation period where practical. Check performance by meaningful segments without creating tiny subgroup claims that cannot be assessed reliably.

Keep calibration and queue quality separate

Guo and colleagues’ calibration research provides a related distinction: a numerical confidence can be unreliable even when a model discriminates among examples. For an account queue, assess both priority usefulness and probability interpretation when probabilities are shown.

A seller should see why an account is selected and what needs validation. A prediction of conversion does not establish intent, identify a named buyer or prove that contacting the account will cause an outcome. Keep those conclusions outside the score.

Explore the original methods and findings in On Calibration of Modern Neural Networks.

Review the mistakes the team cares about

Inspect missed opportunities, irrelevant selected accounts and stale features. Ask whether the failure came from bad labels, missing account relationships or an unsuitable threshold. Review seller overrides with the same outcome window.

Use later cohorts to track changes in event prevalence and workload. The useful report is a capacity-based decision aid with explicit errors, rather than a single accuracy percentage that conceals how the queue behaves.

For the next part of this decision, read Lead-score calibration: when 80 does not mean 80%.

Your next-action checklist

  • Accuracy: Show a confusion matrix. Check the limitation: can hide rare-outcome failure.
  • Precision: Report selected count and population. Check the limitation: changes with prevalence.
  • Recall: Use the actual capacity. Check the limitation: higher coverage can increase workload.
  • Calibration: Compare predicted and observed bands. Check the limitation: different from ranking.

Use the comparison to choose a bounded next step. Record the evidence, the responsible owner, and the review decision before extending the play to additional accounts.

How to use the evidence

Read each reference against the claim it supports. Platform documentation describes capabilities; public cases report a publisher’s experience; research findings apply to the studied task and population. The workflow in this guide is an operating proposal to evaluate in your own account context.

Inspect the research library and connect this guide to measurement and revenue operations.

Questions this guide answers

Why can high prediction accuracy still produce a poor sales queue?

When conversions are uncommon, overall accuracy can hide missed opportunities. Report precision and recall at the team’s capacity, alongside prevalence and the cost of errors.

What should I do first?

Define eligible accounts and outcome. Record the input evidence and the acceptance criteria before continuing. Use the decision worksheet to document the owner, review date and next action.

Read the original research

The guide explains the findings above. Open a publication to inspect its methods, setting and qualifications.

The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. Precision depends on prevalence; comparisons need the same evaluation population.

On Calibration of Modern Neural Networks. These datasets do not validate a particular ABM propensity model.

Connect this guide to the next decision

Lead-score calibration: when 80 does not mean 80% — How can sales tell whether an account score represents a real probability?

How to prioritise ABM accounts when everything looks important — Which accounts deserve scarce research and sales attention this week?

Designing an intent score that explains a prioritization decision — How can an account score help allocate work without claiming to predict a purchase?

PUT IT INTO PRACTICE

Start with your account priorities.

Compare account focus, personalisation, deliverables, and measurement.

Explore Momentum