Give your AI agents access to accurate, real-time customer profiles with FullContact MCP.  No complex setups.

How do you evaluate the quality of an identity graph?

You evaluate the quality of an identity graph by examining three core dimensions: accuracy, coverage, and match rate. A high-quality identity graph correctly links real identifiers to real people, spans a broad range of online and offline signals, and consistently resolves those identifiers at scale. The sections below unpack each of those dimensions in practical terms.

What makes an identity graph high quality?

A high-quality identity graph is one that reliably connects fragmented identifiers (emails, device IDs, cookies, phone numbers, and more) to a single, accurate representation of a real person. Quality is not a single metric but a combination of depth, breadth, freshness, and the integrity of the underlying data sources that feed the graph.

The most important factor is data sourcing. An identity graph built on deterministic signals (data points that are directly provided by the individual, such as a verified email address or a logged-in session) is inherently more trustworthy than one built primarily on probabilistic inference. That said, the best graphs combine both approaches: deterministic anchors provide accuracy, while probabilistic connections extend reach.

Beyond sourcing, freshness matters enormously. People change jobs, move homes, switch devices, and create new accounts. An identity graph that is not continuously updated will degrade over time, producing stale matches that mislead rather than inform. High-quality graphs are refreshed on an ongoing basis, not just periodically.

Depth versus breadth

Depth refers to how many attributes are associated with each identity (professional details, location history, device associations, behavioral signals, and so on). Breadth refers to how many distinct individuals are represented in the graph. Both matter, but they serve different purposes. A graph with enormous breadth but shallow attributes may support audience targeting at scale while falling short for personalization. A graph with rich depth but limited breadth may excel at customer enrichment for known records but struggle with anonymous visitor resolution.

Privacy and compliance as quality signals

A graph built on consent-based, privacy-safe data is not just ethically sound; it is also a stronger quality signal. Data obtained through legitimate means tends to be more accurate because it comes from individuals who actively provided it. Graphs that rely on scraped or opaque data sources carry higher error rates and regulatory risk, both of which undermine quality in practice.

How do you measure the accuracy of an identity graph?

Accuracy in an identity graph is measured by how often the graph correctly links an identifier to the right person and avoids incorrectly merging identifiers that belong to different people. Two specific error types define this: false positives (linking records that should not be linked) and false negatives (failing to link records that should be connected).

The most rigorous way to measure accuracy is through ground-truth testing. This involves taking a known dataset (a set of verified customer records where the correct identity linkages are already established) and running it through the identity graph to see how often the graph’s output matches the verified truth. The closer the output is to the known reality, the more accurate the graph.

In practice, organizations often assess accuracy through several observable proxies:

  • Suppression effectiveness: how reliably the graph prevents duplicate outreach to the same person across channels
  • Enrichment precision: whether appended attributes (such as a job title or location) match what the customer has confirmed
  • Cross-device resolution consistency: whether the same individual is recognized correctly across mobile, desktop, and connected TV environments
  • Downstream campaign performance: whether personalized campaigns built on the graph’s output generate the expected lift

No identity graph achieves perfect accuracy across every use case. The goal is to understand where errors occur and whether the graph’s error profile is acceptable for your specific application. A fraud prevention use case, for example, demands a much lower false-positive tolerance than a general audience segmentation task.

What is a good match rate for an identity graph?

A good match rate for an identity graph depends heavily on the type of identifiers being resolved, the quality of the input data, and the use case. As a general benchmark, match rates for known identifiers like hashed emails against a robust identity graph typically range from 50% to 80%, though this varies significantly based on the data’s age, format, and origin.

Match rate measures how often the graph successfully resolves an incoming identifier to a known identity. A higher match rate means more of your incoming signals (anonymous visitors, new leads, uploaded customer lists) can be tied to an existing profile and enriched with additional attributes. But match rate should never be evaluated in isolation.

Match rate versus match quality

A graph that inflates its match rate by making loose or probabilistic connections may report impressive numbers while delivering inaccurate results. This is why match rate and accuracy must be evaluated together. A 75% match rate with high accuracy is far more valuable than a 90% match rate where a significant portion of matches are incorrect or unreliable.

Factors that influence match rate

Several variables affect how high a match rate you can realistically expect:

  • Identifier type: hashed emails and phone numbers typically yield higher match rates than cookies or device IDs, which are more volatile
  • Data recency: older customer lists with outdated emails or inactive accounts will naturally produce lower match rates
  • Graph coverage: the size and diversity of the identity graph’s underlying population determine how many of your inputs can be resolved
  • Industry vertical: B2B datasets often see different match rate profiles than B2C datasets due to the nature of the identifiers involved

When evaluating a vendor’s match rate claims, always ask for results against data that mirrors your own in terms of recency, identifier type, and industry. A benchmark derived from a different data type or audience profile may not reflect what you will actually experience.

How FullContact helps you evaluate and apply a high-quality identity graph

We built our Resolve platform specifically to address the quality dimensions that matter most when evaluating an identity graph. Here is how we approach it in practice:

  • Real-time resolution at scale: our identity graph returns API responses in under 150 milliseconds, resolving authenticated and anonymous identifiers into a single customer profile without latency
  • 900+ enrichment attributes: each resolved identity can be appended with personal and professional insights, giving you depth alongside coverage
  • Privacy-safe by design: our graph is built on consent-based data, which means the accuracy and compliance standards are built into the foundation, not added as an afterthought

Whether you are trying to improve match rates on an existing customer list, resolve anonymous visitors in real time, or enrich new leads before they enter your CRM, our identity graph is designed to deliver reliable, accurate results across all three quality dimensions. If you want to see how our graph performs against your specific data and use case, contact us and we will walk you through it.

Related Articles

What Can We

Create Together?