Data quality

9 min read

How we decide what to return, what each response tells you about its own reliability, and how to trade match rate against certainty.

Bad data costs more than missing data. A wrong profile written to your CRM gets copied, reported on, and emailed before anyone notices. A missing profile is just a gap you can fill later.

That asymmetry shapes the whole API. This guide explains the principle behind it, the signals every response carries about its own reliability, and how to tune the balance between match rate and certainty for your use case.

A not-found is a real answer

When we cannot identify a person or a company with confidence, we say so. You get a 404, or a documented error code on the async endpoints, instead of a plausible-looking record assembled to fill the response.

That is a deliberate design choice, and it runs through the whole platform: filters that we cannot honour are rejected instead of silently dropped, a job that failed reports why, and a profile removed by its owner stays removed. Every one of those cases is a place where we could have returned something instead of nothing. We would rather hand you a clean gap you can act on than a record you have no way to distrust.

Three questions behind every response

Every record you receive answers three separate questions, and each one has its own signal in the payload.

  1. Is this the right person? Identity resolution.
  2. Is this still true? Freshness.
  3. Can I act on it? Deliverability.

A response can score well on one and poorly on another. A profile can be the right person and eighteen months out of date. An email address can be perfectly formed and belong to someone who left the company last year. Read the three signals separately.

1. Is this the right person?

Two ways to identify someone

The API has two distinct families, and they answer to different inputs.

Family Endpoints What you send
Fetch Profile, Profile status, Profile live A LinkedIn profile URL, or a Reverse Contact id (prs_...)
Enrich Enrich profile, Enrich company Attributes: email, firstName, lastName, companyDomain, companyName, or a company domain

Fetch assumes you already know which profile you want and gives you its content. Enrich is the one doing identity resolution: it takes the attributes you have and decides who they point to. Everything below about ambiguity applies to Enrich.

Find Email sits in between: it accepts a LinkedIn profile URL, or a full name with a company domain.

Match rate and certainty pull in opposite directions

Every field you add to an Enrich request narrows the search. That is a trade-off, not a free win.

What you send Match rate Certainty it is your person
firstName + lastName Highest. Almost always returns someone Lowest. There are hundreds of people with a common name
firstName + lastName + companyName High. Company labels are written many ways Moderate. Company names are not unique
firstName + lastName + companyDomain Lower. Only matches if that person is known at that domain High
email Lower. Only matches if we know that address Highest. One address points to one person at one company

Neither end of that table is the correct answer. It depends on what a wrong match costs you:

  • Automatic CRM writes, outbound sequences, anything a customer sees. Prefer certainty. Send companyDomain or the email, accept more 404s. A wrong profile is worse than no profile.
  • Coverage work reviewed by a human, or a first pass you intend to refine. A looser query is fine. Get the candidates, then decide.

The same logic applies to Find Email: a name-and-domain lookup can resolve to the wrong person, and then the address it discovers is a correct address for the wrong human being.

A LinkedIn URL is not a permanent identifier

It is tempting to treat a profile URL as a primary key. It is not one. A public slug is chosen by the person who owns the profile, and they can change it whenever they like. When they do, the URL you stored six months ago points at nothing, and your call comes back as 404, or as 422 INVALID_LINKEDIN_URL if the stored value is malformed.

Store our identifiers instead:

  • id (prs_... for a person, com_... for a company) is stable across slug changes. Every response that returns a person or a company carries it.
  • Both Profile and Profile status accept id in place of url, and id wins when you send both.
  • Profile live resolves the id to the current public slug before it starts, so a live refresh keeps working after a rename.

Keep the URL for display and for humans. Key your records on id.

Use alternativePersons as an ambiguity flag

When several people match your input, Enrich profile returns the best match in data.person and the other candidates in data.alternativePersons, ordered best first, up to nine.

Treat that array as the confidence signal it is:

  • Empty. One person fits what you sent. Safe to write automatically.
  • Not empty. Several people fit. Add an identifying field and retry, or route the record to human review before it reaches a CRM or a sequence.
Ambiguity gate
JavaScript
const { data } = await enrichPerson({ firstName, lastName, companyDomain });

if (data.alternativePersons.length > 0) {
  // Several people match this input. Do not write automatically.
  await queueForReview(data.person, data.alternativePersons);
} else {
  await writeToCrm(data.person);
}

Two details worth knowing before you build on the alternatives: each one is a full profile, the same shape as data.person, so you can review candidates side by side without an extra Profile call, and their currentCompanyId is always null, so you cannot chain a company fetch directly from a candidate you have not confirmed yet.

Current position is a decision, not a raw field

A profile can list several jobs with no end date: a main role, a board seat, an advisory gig. currentPosition is the one we single out, and we pick it with fixed, published tie-breakers rather than a heuristic that drifts between calls. The experience and education arrays are ordered by the same fixed rules, whether the record came from our database or from a live fetch.

So when a title looks wrong, the cause is almost always a missing or outdated date in the source profile, not an unpredictable API. Current position and work history ordering walks through exactly how the choice is made.

2. Is this still true?

Professional data decays continuously. People change jobs, companies rename, teams get restructured. A record that was accurate when we captured it is not accurate forever, so every record tells you when it was last refreshed.

Where you look Field What it means
Enrich profile, People search updateDate When we last refreshed that record
Profile, Company profile metadata.updatedAt Last upstream update, present only when the response was served from cache
Profile status, Company status lastUpdate When we last fetched the profile from the source, null when it does not exist

Three ways to act on freshness instead of hoping for the best:

  • Read it before you spend. Profile status costs 0 credits and returns exists plus lastUpdate, along with experienceCount, skillCount and schoolCount so you can spot thin records too.
  • Filter it at the source. People search and Company search accept maxDataAgeDate, an ISO 8601 datetime that returns only profiles refreshed after that date. Prefer it over filtering stale rows in your own code after you paid for them.
  • Refresh on demand. The live endpoints pull a fresh copy when the stored record is too old for your use case.

Pick an explicit threshold per use case instead of treating every record the same:

Use case Suggested threshold If older
Outbound sequence, meeting prep 30 days Refresh live before use
CRM enrichment at scale 90 days Refresh in your next batch
Analytics, market sizing 180 days Usually fine as is

Data freshness covers the Check → Fetch → Live pattern and what it saves.

3. Can I act on it?

Find Email returns one entry per address discovered, each with a value and a type, typically professional. It is asynchronous: you get a webhookId, then retrieve the result by polling or through a webhook callback. When nothing is found you receive email-not-found rather than a constructed address.

Two habits protect your sender reputation whatever the source of an address:

  • Keep your own suppression list authoritative. Unsubscribes, complaints and hard bounces you have already recorded always win over a freshly discovered address.
  • Track bounce rate per input type. If one segment bounces above your baseline, the input that produced it is usually the problem rather than the addresses themselves. Name-based lookups and URL-based lookups deserve separate counters, for the reason described earlier.

Note

An address that exists is not the same as an address you should contact. Consent and applicable regulations remain yours to handle: the API tells you what exists, not whom you may email.

Filters that fail loudly

Search deserves its own note, because a bad filter is the quietest way to get bad data. If we ignored a field we did not recognise, you would receive a full page of results drawn from a query you never asked for, and nothing in the response would look wrong.

So the search endpoints validate every filter before the query runs. A rejected request returns 422, lists every problem in error.details.errors, and costs 0 credits. Unknown field names, blank values, values outside a closed list such as currentCompanyIndustry, and inverted ranges are all refused rather than dropped. See Search filter validation for the complete list.

Two behaviours to keep in mind when you read search results:

  • Title matching is token-based, not exact. "CTO" also matches "Deputy CTO". Multiple titles are OR-matched, and excludeCurrentPositionTitle always wins over a match.
  • Headcount is stored as fixed brackets. A range only matches a bracket it fully contains, so bounds that fall inside a bracket return an empty array, and that request is still billed. Send bracket boundaries.

When a profile is blocked

A profile removed under a GDPR or CCPA data subject request returns 451. That state is permanent: the record is not stale, not missing, and will not come back. Retrying wastes calls, so treat 451 as a terminal outcome in your pipeline and drop the record from your enrichment queue rather than leaving it to be picked up on the next pass.

Putting it together

For a pipeline that writes to a system of record, this is the shape we recommend:

Decision tree
Text
For each lead:
  1. Enrich with the most identifying fields you have
     │
     ├─ 404 (no confident match)
     │  └─ Retry with fewer fields → route the result to review
     │
     └─ Match
        │
        ├─ alternativePersons not empty?
        │  └─ Route to human review, do not write
        │
        └─ Unique match
           │
           ├─ Store the id (prs_...) and updateDate
           │
           └─ updateDate outside your threshold?
              └─ Refresh live by id, then write

The two gates that matter are the ambiguity gate and the freshness gate. Skipping either is how wrong data enters a CRM quietly.

Quality checklist

  • Use Fetch when you already know the profile, and Enrich when you are still resolving who it is.
  • Key your records on our id, not on a LinkedIn URL that its owner can rename.
  • Never write a record with a non-empty alternativePersons array without review.
  • Choose your input mix deliberately: certainty for automated writes, coverage for reviewed work.
  • Store the freshness timestamp next to every record so you can audit staleness later.
  • Set a staleness threshold per use case, filter with maxDataAgeDate on search, and refresh live above it.
  • Treat 404 as a normal outcome to log and 451 as terminal, not as errors to retry blindly.
  • Keep your suppression list ahead of any discovered address.
  • Track match rate and bounce rate per input type. They tell you which inputs to improve.

Previous

Which endpoint should I use?

Next

Error handling & retries