Methodology, protocol v1

How we benchmark B2B data providers

Six steps, the same for every provider and fixed before the run they govern, ours included. If a number on any report disagrees with this page, this page is the one that is wrong and we will fix it.

1. The run, step by step

One run is one photograph: the same list, the same input, the same rules. Here is the order in which it happens, and nothing happens outside it.

  1. 01

    The list

    A run starts from lists of LinkedIn profile URLs, each list isolating one situation: people who changed jobs recently, leaders in companies of a given size, a real CRM export. Each list is tagged as a cohort, and the tag follows every contact into every table, so any figure on a report can be broken back down into the populations that produced it.

  2. 02

    Who is left out, before anyone is queried

    Contacts are removed before any provider sees the list, never after the answers are in. In the latest run the causes were: people who started the current job in the month of the test, so no index could reasonably have seen it yet, have no current role on their profile, so there is nothing to be right or wrong about, have fewer than 30 connections, our floor for an active account and have no start date on the current role, so job age cannot be computed. Every report publishes how many contacts were removed under each cause.

  3. 03

    A reference that belongs to nobody

    On the day of the test, the current employer and title of every contact are read from their public profile, from a source that is none of the providers tested. Only fields any visitor can see are read, no contact details enter the reference, and LinkedIn is not involved in these tests. A run only proceeds when the reference is essentially complete. Measured on September 17, 2026 on 3,445 profiles, the latest run read 3,916 public profiles first, 100% of those requested, and the filter above brought that down to the 3,445 contacts every provider was then sent.

  4. 04

    Same list, same input

    Every provider receives the identical list, one call per contact, through its own public API, with the same input: the person's LinkedIn profile URL. Transport errors are retried until the run is clean. A contact a provider genuinely does not have counts against coverage, not against accuracy.

  5. 05

    One verdict per answer

    Employers are matched by identifier first, then by URL, then by normalized name, because two companies can share a name and one company can be written five ways. When the employer matches, the title is compared after normalization. Each returned contact ends in exactly one of four verdicts; a contact a provider never returned is in none of them and counts against coverage only.

  6. 06

    Recomputed, then published whole

    Before anything goes online, a separate script reads the raw answers again, recomputes every verdict and every aggregate, and checks that they add up. Its result is published per provider, warnings included: when a provider passes with a warning, the report says which check and what it affects.

2. Why it is built this way

Every step above closes off a way of getting a flattering number. These are the six questions worth asking of any benchmark, including this one.

Why not score against a provider’s own database?

A benchmark scored against one vendor’s database measures similarity to that vendor, not accuracy. The reference has to come from outside all of them, or the result is a measure of agreement.

Why give every provider the same input, instead of letting each one do what it does best?

Every provider is handed the person: the profile URL, which all of them accept. The test measures the record behind a known person, not the ability to find that person from a name and a company. Mixing the two would measure two things at once and hide which one gave way.

Why remove contacts before the answers come in?

Excluding contacts after seeing the answers is the easiest way to shape a benchmark, and it is the first thing worth checking in anyone else’s. The filter runs before the queries, and the count removed under each cause is published.

Why is there no price column?

None of the providers tested publishes a per-unit price we can verify and date. Rather than compare quoted rates, the reports publish no price and no billing unit at all: ask each provider for the rate it quotes you, and apply it to the volume you send.

Why is nothing ranked when the gap is small?

Three rates carry a 95% Wilson confidence interval, because they are proportions of a known base: coverage, accuracy, and accurate records per 100. A segment under 30 contacts is not published at all, one under 100 is published but not ranked, and a pair whose gap is smaller than the combined margin is marked too close to separate, even when the sort still puts one above the other. Every other measure prints its gap and says the separation was not tested, rather than calling a lead it has not verified.

Why publish a test we are in?

We sell one of the products measured. The reference comes from none of the providers, the rules were fixed before the run, the recomputation is separate, and the segments where our own product falls behind are published in the same tables as everyone else’s. That reduces the conflict; it does not remove it. Measure us on your own list before believing any of it.

3. The words on the figures

A technical reader loses a report on a word, not on an idea. Seven words carry every figure we publish.

Coverage

The share of the contacts we queried for which the provider returned a person at all.

Contacts returned divided by contacts queried. A returned contact counts here even when its data is wrong.

Accuracy

Among the contacts a provider did return, the share where it names the right current employer and a matching title.

Right employer and title divided by contacts returned. Right employer means the employer the person's public profile showed as current on the day of the test.

Accurate records per 100 contacts

Out of 100 contacts you send, how many come back usable. Coverage multiplied by accuracy.

Right employer and title divided by contacts queried, times 100. This is the ranking measure, because it is the only one a provider cannot improve by answering less often or by answering more loosely.

Stale employer

The provider names an employer the person has left.

Split into lag, where the named company really appears earlier in the person's history, and never worked there, where it does not appear at all.

Unverifiable

Neither side carries enough data to judge the answer.

No usable company on the provider side, or no current role in the reference. Published so the four verdicts always add up to the contacts returned.

Work history recall

How much of a person’s career the provider knows, not just the current job.

Reference experience rows matched in the provider answer, divided by reference experience rows, over the profiles the provider returned. Read it beside the number of roles an endpoint returns at all: an endpoint built to return one role scores low on scope, not on quality.

Too close to separate

The gap between two values is smaller than the uncertainty on them, so the test does not call one ahead of the other.

The two 95% confidence intervals overlap. The sort still puts one above the other, and the page says so rather than hiding it: a tie is a result, not a missing value.

The four verdicts

The four verdicts
VerdictWhat it meansBase
Right employer and titleThe current employer matches the reference, and the title on that role matches after normalization.Contacts returned
Right employer, different titleThe current employer matches, the title does not, or is missing on one side.Contacts returned
Stale employerThe current employer does not match the reference. Usually a real previous role, sometimes a company absent from the history entirely.Contacts returned
UnverifiableNo usable company data on one side, or no current role in the reference.Contacts returned

A contact a provider never returned is not in these four: it counts against coverage only. The four verdicts always add up to the contacts returned, which is why the shares on a report are shares of returned contacts and say so.

4. What this test does not tell you

The reference is one day old, by construction

A provider that indexes a change the day after the test is measured as wrong, and it is. That is why accuracy is published by job age rather than as a single number, and why a run is repeated rather than cited forever.

The population is chosen, and it is not the market

Cohorts are drawn to isolate the situations that break contact data. A benchmark built on established contacts only would produce much higher numbers for everyone.

The test measures the record, not the search

A provider whose strength is finding someone from a name and a company is not measured on that strength here. Every provider is handed the person.

Not every provider was queried on the same day

One extra day is shown for each provider in the run conditions of every report, and it works in their favor, not against them.

A title match tolerates wording

Titles are compared after normalization, so a rewording counts as a match and a real promotion does not. The direction of title mismatches is measured and published.

5. What we publish, and how to get a figure corrected

Every figure is published on the report page itself, and the same report is downloadable as a PDF, free and without a form. The PDF is the page, printed, and every figure in it comes from the same data. No row in either identifies a person.

Row level files are a different matter. Each row describes a real professional profile and carries a judgement about their current job, so they are not published. The full rule set behind a run is sent to any provider in it, on request, with the test id.

A person who appears in a reference can ask to be removed from future runs by writing to antoine@blitz-api.ai.

Corrections and right of reply

Every named provider receives its results before publication, with the method and the verdicts that concern it, and can reply. Replies are published on the report as sent, and a provider that does not reply is recorded as not having replied, without comment.

  • Corrected: a wrong value, a calculation error, a billing rule we applied wrongly, a rule described here that the code does not implement. Corrections are versioned and stay visible.
  • Published as a reply rather than a correction: disagreement about whether the measure is the right one, or about the population. The objection goes next to the number it contests.
  • What we will not do: quietly edit a published figure, remove a cohort after seeing the results, or delay a run because it came out badly for us.

Write to antoine@blitz-api.ai with the test id of the run. We acknowledge within two business days and correct or retract within five.

Editions, and the protocol each one ran on

A new report roughly every month, when the run separates providers and the protocol has not moved mid-flight. The version number moves only when a change would alter a figure already published, if the earlier run were scored again.

Run the same test
on your own list

Bring your own reference and measure any provider, including us.