Blog/Sales & enrichment
12 min read

AI Lead Score Calculation: The Inputs Disagree Before the Model Runs

Two providers gave one company 580 and 1,011 employees. A third resolved a famous domain to a different company with a plausible record.

AI Lead Score Calculation: The Inputs Disagree Before the Model Runs

Copy this line to your agent to build a scoreable record for one domain.

set up https://monid.ai/SKILL.md and use apollo /organizations/enrich plus hunterio /email-count

We scored three company domains on 2026-09-23 and the interesting part happened before any scoring logic ran. Two providers gave one company 580 and 1,011 employees. One provider resolved a domain everybody knows to a completely different company, and returned a tidy, plausible, internally consistent record for it. A lead score is arithmetic over fields, and the fields do not agree with each other. This guide runs through Monid, the OpenRouter for agent tools.

What goes into an AI lead score?

Four kinds of input, and only one of them is about the lead.

Fit

Firmographics: size, industry, geography, age, funding. This is the bulk of most scores and it is what an enrichment call returns. It answers "is this the kind of company we sell to".

Capability

Technology, headcount by function, tooling. This answers "could they actually use us", which is different from whether they match the profile. The measurement below produces a capability signal from a field nobody thinks of as technographic.

Reachability

How many contacts exist, in which departments, and whether you can get an address at all. A perfect-fit company you cannot reach scores zero in practice and 90 on most models.

Intent

Behaviour: visits, downloads, job postings, funding events. Comes from your own systems and from event sources, not from a firmographic call, and it decays in days rather than quarters.

Where the model sits

"AI lead score" usually means a model weighting those inputs, and the weights are the part people argue about. The measurement below is about the part nobody checks: whether the inputs are true.

📖 See also Target Company URL Analysis: What One Domain Tells You in Four Calls

Why do two providers disagree by 74 percent?

Because headcount is modelled, not counted, and each provider models it from a different population.

One domain, two answers

vercel.com, 2026-09-23:

FieldApolloHunter
Employees5801,011, band 1K-5K
Founded20152015
Moneyannual_revenue 340,000,000raised 863,000,000
Technologies18746

Founded year agrees. Headcount differs by 74 percent. And the money fields are not two estimates of one quantity at all: one is annual revenue, the other is capital raised. A scoring model that maps "the money field" to a band will cheerfully treat 863 million of venture funding as revenue.

What to do about it

Pick one provider per field and write it down. Not one provider per record. Headcount from whichever you trust on headcount, funding from whichever reports funding, and never fall back to the other one when a value is missing, because a backfilled field silently changes the units of your score.

Store the provider on every value. Six months later the only way to explain why a company moved from band B to band A is knowing which source answered that day.

Band, do not thresh. 580 and 1,011 fall in the same band on any sensible scale and on opposite sides of a threshold at 1,000. Thresholds convert provider disagreement into score volatility.

Treat the technology count as a source property. 187 against 46 is not a measurement difference, it is a definition difference, and the three-way version of that gap is measured in the technographics guide.

How do you calculate a lead score through one key?

Four steps, and the third one is the one that saves you from the next section.

For agents

Grab an API key at app.monid.ai, then paste this to your agent and hand it the key:

set up https://monid.ai/SKILL.md

It learns the whole discover, inspect, run workflow itself. More in the agent quickstart.

For humans

npm install -g @monid-ai/cli
monid keys add -k <your-api-key> -l main

Step 1. Pull fit, one call per domain

monid run -p apollo -e /organizations/enrich --query '{"domain": "vercel.com"}'

Sixty-six fields including industry, founding year, headcount, revenue estimate and a technology list. Two to three seconds per domain on our runs.

Step 2. Pull reachability, one more call

monid run -p hunterio -e /email-count --query '{"domain": "vercel.com"}'

Returns total known addresses split into personal and generic, plus a department breakdown. This is the cheapest useful signal in the whole exercise and most scoring models never ask for it.

Step 3. Cross-check before you score

Compare the two records. If the headcount says three people and the address count says fifty-three across eight departments, one of them is wrong and the next section is what that looks like.

suspect = org["estimated_num_employees"] < 10 and counts["total"] > 40

One line. It would have caught our bad row.

Step 4. Score on bands and ratios, not raw values

fit       = band(headcount) + industry_match + age_band
reach     = band(counts["total"]) + has_department(counts, "executive")
capability= counts["department"]["it"] / counts["total"]
score     = w1*fit + w2*reach + w3*capability

The third line is free and it is the subject of the last part of this guide.

Give this to your agent

$Set up https://monid.ai/SKILL.md, and then use Monid to for each domain in this list, pull the company record and the email count, flag any row where the employee estimate is under ten but the address count is over forty, and give me the IT share of known contacts for the rest.

📖 See also Turn a Domain Into Full Company Firmographics in One Call

Why did a famous domain score three out of a hundred?

Because the provider resolved it to a different company, and the record it returned was perfectly sensible for the company it found.

What came back for basecamp.com

{ "name": "Basecamp Talent",
  "industry": "staffing & recruiting",
  "linkedin_url": ".../company/basecamp-talent",
  "founded_year": 2023,
  "estimated_num_employees": 3,
  "annual_revenue": 12300000,
  "short_description": "…a UK-based recruitment agency…" }

A three-person recruitment agency founded in 2023. Every field is plausible, the numbers are consistent with each other, and the LinkedIn URL points at a real company. Nothing about this record looks broken.

It is attached to the domain of a well-known software company that has existed since the 1990s.

Why this is worse than a bad estimate

A wrong number gets noticed. A wrong company does not. If the headcount had come back as 3 for the right company, a human reviewing the list would flag it. A coherent record for a staffing agency passes every sanity check you would write about internal consistency, because it is internally consistent.

It poisons the score in both directions. Our software company was disqualified by an industry filter looking for technology, by a headcount floor, and by a founding-year rule about maturity. Three independent parts of the model agreed, and all three were scoring the wrong company.

And it is silent. No error, no null, no confidence field saying otherwise. The same silent-failure class as the derived company domain in the H-1B guide, where a plausible-looking string was a join key to nothing.

The cheap defence

Hunter's count for the same domain: 53 known addresses, 24 personal and 29 generic, spread across support, IT, executive, finance and HR. A three-person agency does not have that. One extra call, a fraction of a cent, and the contradiction is visible.

Cross-check on a quantity, not a name. Comparing two providers' company names would not have helped, since both would say "Basecamp". Comparing headcount against address count catches it immediately, because the two numbers have to be roughly compatible and here they were not.

Which endpoint should I use for which job?

EndpointWhat it doesInputOutputBest forBilling
apollo/organizations/enrichCompany recorddomain66 fields: industry, headcount, revenue, technologiesThe fit inputsPer call
hunterio/email-countAddress counts by departmentdomainTotal, personal, generic, per-departmentReachability and the cross-checkPer call
hunterio/companies/findFirmographics, second opiniondomain34 fields including a headcount bandResolving a disagreementPer call
api.strale.io/x402/tech-stack-detectLive technology detectionurlNamed technologies with evidenceVerifying one capability claimPer call
hunterio/multi-domain-searchPopulation with existence flagsFiltersRedacted rows with flagsSizing a segment before scoring itFree

Every row was verified with monid inspect on 2026-09-23. The table gives billing shape rather than figures, because shape drives design and current numbers live on monid.ai/tools.

The second row is the cheap one that makes the first row trustworthy, which is the opposite of how most scoring stacks are built.

When is a score the wrong tool?

Four cases.

Your list is small. Under a few hundred accounts, a person reading them beats a model, and the model's errors cost more than the reading time.

Your criteria are binary. If you only sell to companies in three industries above a headcount floor, that is a filter, not a score. Collapsing a filter into a weighted number hides which rule rejected a company.

You have no intent data. Fit and reachability describe a company's suitability, not its readiness. A score built only from firmographics ranks the market, not the pipeline, and will be stable for months because the underlying fields barely move.

The entity resolution is not solved. Everything above is downstream of knowing which company a domain belongs to. If your inputs include a staffing agency wearing a software company's domain, improving the weights is wasted work. Fix resolution first, which is what the company lookup guide is about.

And the disclosure: this is Monid's blog and we resell every endpoint here. The central finding is that one of them attached the wrong company to a well-known domain, and the fix is to spend a fraction of a cent on a second call to catch it.

Conclusion

An AI lead score is arithmetic, and arithmetic inherits its inputs. On 2026-09-23 two providers described one company as having 580 and 1,011 employees, reported two different quantities in similar-looking money fields, and counted its technologies as 187 and 46. On a second domain, one provider returned a coherent, plausible, entirely wrong company: a three-person recruitment agency founded in 2023 sitting on a software company's domain, which three separate parts of a scoring model would independently reject.

What matters more than the weights is the plumbing. Pick one provider per field and never backfill from the other. Band rather than threshold, because 580 and 1,011 belong in the same bucket and sit on opposite sides of a cutoff. Add one cheap call that counts addresses, both because reachability belongs in the score and because it is the contradiction that catches a wrong entity. And look at the department mix: IT was 34 percent of known contacts at a software company and 6 percent at a retailer, which is a capability signal you get for free.

Free next step: run both calls on five domains you know well and compare the headcounts. If any pair differs by more than a band, you have just learned something about your score that no amount of model tuning would have told you. Start at monid.ai.

FAQ

What does a lead score of 95 actually mean?

Whatever your weights say it means, which is why the number travels badly between teams. A 95 is only interpretable alongside the band definitions and the field sources behind it, and the same company can score 95 and 60 under two models built from the same data. Publish the components rather than the total: fit band, reachability band, capability ratio, and the provider that supplied each input. A rep can act on "large, reachable, technical, and the headcount came from one source we have not cross-checked". Nobody can act on 95.

Should you use one provider or several?

Several, but one per field, decided in advance and written down. Our two providers agreed on a founding year and differed by 74 percent on headcount for the same domain, and they reported different quantities entirely in their money fields, so blending them produces a number with no units. The pattern that works is a primary source per attribute plus one cheap secondary call used only as a contradiction check, never as a backfill. Backfilling is what turns a consistent score into one that quietly changes meaning depending on which provider happened to have the row.

Can you score capability without a technographic call?

Yes, and it is more robust than the technology lists. The department breakdown on an address-count call gave us IT as 34 percent of known contacts at a software company against 6 percent at a clothing retailer, from one cheap call with no technology data involved. That ratio is hard to game, it does not depend on whose definition of "uses" you accept, and it does not suffer the 271-versus-18-versus-zero problem measured in the technographics guide. Use technology lists to confirm a specific claim, and use the department mix to characterise the company.

How often should you recompute a score?

Fit inputs move on a quarterly timescale, so recomputing firmographics weekly spends money to observe noise. Reachability moves faster as people join and leave, and intent moves in days. The practical split is to refresh firmographics quarterly or on a trigger such as a funding event, refresh address counts monthly, and recompute the intent component as events arrive. What matters more than the cadence is storing the fetch date with every value, because a score assembled from a fresh reachability number and a year-old headcount is not wrong so much as undated, and undated is harder to debug.

Last updated September 2026.

ai lead score calculation methodlead scoringicp scoringlead qualificationfirmographics