Scoring leads with conformal abstention
A model that confidently ranks a bad lead high is worse than one that says "I am not sure." The most useful property a scoring system can have is the discipline to abstain.
Every lead-scoring pitch sounds the same: point our model at your pipeline and it will tell you which leads to call first. The unspoken assumption is that a score is always trustworthy. It is not. A model trained on your history will be confident about the leads that look like your past wins and quietly guess about everything else. The dangerous outputs are not the low scores — those get ignored, which is fine. The dangerous outputs are the confident-but-wrong high scores, because a rep will spend real hours on them.
We think the most valuable property a scoring system can have is not accuracy. It is the discipline to say "I am not sure."
This is the idea behind conformal prediction, a methodology worth understanding even if you never implement it yourself. The ordinary way to score a lead is to emit a single number — 0.82, call this one. Conformal methods do something different. Instead of a point estimate, they produce a set of plausible outcomes calibrated to a confidence level you choose. On an easy, familiar lead that set is small and sharp: this will convert. On an ambiguous lead the set is wide, and the honest reading of a wide set is abstention — the model is telling you it does not have enough signal to rank this one, and you should not treat its guess as information.
Abstention feels like a weakness. It is the opposite. A model that abstains on the 15% of leads it genuinely cannot judge, and commits only on the 85% it can, gives a rep a cleaner instrument than a model that assigns a crisp-looking score to all of them. False confidence is expensive precisely because it is invisible: nothing on the screen distinguishes a well-grounded 0.82 from a hallucinated one. Calibrated abstention makes that distinction visible, which is the entire point.
Now the honest part, because we would rather be trusted than impressive.
Propel does not ship a conformal predictor today. What we ship is a transparent, heuristic model. Deal health and lead signal come from computeDealHealth — a set of readable rules over the things that actually correlate with momentum: recency of contact, stage progression, engagement, time decay. We chose a heuristic on purpose for this stage of the product. A heuristic has a property a black-box model does not: a rep can look at why a deal is flagged and either agree or overrule it. When you cannot yet calibrate a statistical model honestly, an explainable rule you can argue with beats a confident number nobody can audit.
Machine-learned scoring is on our roadmap, and it will run in the Python side-car service that sits behind our edge — the same place we are building embeddings and search. But we are treating the ML version as a promise we have not yet kept, not a feature we quietly imply. When it ships, the design principle above is the bar it has to clear. A learned scorer that cannot abstain, that cannot tell us on which leads it is out of its depth, is not an upgrade over an honest heuristic. It is just a more persuasive way to be wrong.
There is a broader lesson here for anyone shipping AI into a workflow that costs real time. The instinct is to maximize how often the model gives an answer. The better instinct is to maximize how often the model gives an answer you can trust, and to make the remainder — the "I don't know" — a first-class, visible output rather than a confident shrug.
A scoring model's willingness to abstain is not a gap in the product. It is the feature. Any tool can hand you a ranked list. The one worth trusting is the one that tells you where the ranking stops being reliable.