An identity graph links identifiers by connecting all the known data points associated with a real individual (email addresses, device IDs, cookies, phone numbers, and more) into a single, unified profile. Rather than storing these identifiers in isolation, an identity graph maps the relationships between them, so that when one identifier is recognized, the full picture of that person becomes accessible. The sections below unpack what types of identifiers are stored, how connections are made, and what separates deterministic from probabilistic matching.
How are different identifiers connected inside an identity graph?
Identifiers inside an identity graph are connected by mapping overlapping data points across multiple sources and interactions to a single individual. When two or more identifiers appear together consistently (for example, the same email address being used to log in on a mobile device and a desktop browser), the graph creates a persistent link between them, building a web of relationships that represents one real person rather than a collection of disconnected signals.
The connection process works by continuously ingesting identity signals from online and offline touchpoints. Each time a person interacts with a brand (visiting a website, opening an email, making an in-store purchase, or using a mobile app), they leave behind identifiers. The identity graph captures these signals and evaluates whether they belong to an existing profile or represent a new one.
Several mechanisms drive how identifiers get joined together:
- Authenticated events anchor the graph. When a user logs in or provides an email address, that known identifier becomes a reliable tie point for linking anonymous signals.
- Co-occurrence patterns reveal relationships. If the same cookie and the same device ID appear repeatedly in the same session, the graph treats them as belonging to the same person.
- Third-party data partnerships extend the graph’s reach, filling gaps where first-party data alone cannot establish a connection.
What makes an identity graph fundamentally different from a standard customer database is the graph structure itself. A relational database stores rows of records. An identity graph stores nodes (the identifiers) and edges (the relationships between them), allowing it to traverse connections and resolve identity even when only one identifier from a cluster is present. This architecture means the graph can recognize a returning visitor across devices without requiring them to log in every time.
What types of identifiers does an identity graph store?
An identity graph stores a broad range of identifiers that span both online and offline data. These typically include personally identifiable information such as email addresses and phone numbers, digital identifiers such as cookies, mobile advertising IDs, and IP addresses, as well as offline signals like postal addresses and loyalty program IDs. Together, these data points form the raw material the graph uses to build and maintain individual profiles.
It helps to think about identifiers in two broad categories: persistent and transient.
Persistent identifiers
Persistent identifiers tend to stay stable over time and are closely tied to a real individual. They form the most reliable anchors inside an identity graph because they are unlikely to change frequently or be shared between multiple people. Examples include:
- Email addresses (especially hashed versions used in privacy-safe matching)
- Phone numbers
- Home or mailing addresses
- Loyalty card or account IDs
Because persistent identifiers carry a high degree of confidence, they are often used as the primary keys around which other, more transient identifiers are clustered.
Transient identifiers
Transient identifiers are digital signals that can change or expire. Cookies reset when a browser is cleared. Mobile advertising IDs can be reset by the user. IP addresses shift when someone moves between networks. Despite their impermanence, these identifiers are valuable because they capture real-time behavioral signals and help the graph extend identity into anonymous or semi-anonymous environments where persistent identifiers are not available.
Beyond these two categories, modern identity graphs also incorporate professional and contextual data (job titles, company affiliations, social handles, and demographic attributes) that enrich individual profiles without necessarily serving as direct linking keys. These enrichment attributes add depth to what the graph knows about a person, enabling more meaningful personalization once the identity connection has been established.
What’s the difference between deterministic and probabilistic identifiers?
The key difference is certainty. Deterministic identifiers create identity matches based on confirmed, directly observed data: a user logs in with their email address, and that link is known with high confidence. Probabilistic identifiers create matches based on statistical inference: patterns, signals, and behavioral data suggest that two identifiers likely belong to the same person, but without absolute certainty.
Both approaches play important roles inside an identity graph, and most sophisticated graphs use them in combination rather than relying exclusively on one.
Deterministic matching
Deterministic matching occurs when a user provides a known identifier directly. A login event, a form submission, or a purchase confirmation all create a clear, auditable link between a person and a digital touchpoint. Because the match is based on first-hand data, deterministic connections carry the highest confidence score inside the graph. The trade-off is reach: deterministic data is only available when a user actively identifies themselves, which covers a fraction of total digital interactions.
Probabilistic matching
Probabilistic matching fills the gaps left by deterministic data. It uses signals such as shared IP addresses, device type, browsing patterns, and location data to infer that multiple identifiers belong to the same individual. The confidence level varies depending on how many corroborating signals align, and the graph typically assigns a probability score to each inferred connection. High-quality probabilistic matching can dramatically extend the reach of an identity graph, allowing brands to recognize and engage individuals across a much larger share of their digital footprint.
The practical implication for marketers and data teams is that neither method alone is sufficient. Relying only on deterministic data limits the graph’s coverage. Relying only on probabilistic data introduces more uncertainty into targeting and personalization decisions. The strongest identity graphs layer both, using deterministic anchors as a foundation and probabilistic inference to extend coverage where direct signals are absent.
How FullContact helps with identity graph resolution
We built our Resolve platform specifically to address the complexity of linking identifiers at scale, in real time, and in a privacy-safe way. Our identity graph has been developed over more than a decade and encompasses both online and offline data, all organized around real individuals rather than devices or sessions. Here is what that means in practice:
- Real-time resolution in under 150 milliseconds, so identity matching happens at the moment of interaction rather than in batch processing after the fact.
- Support for authenticated and anonymous identifiers, combining deterministic and probabilistic matching to maximize both accuracy and reach.
- Enrichment with 900+ personal and professional attributes, appended to unified customer profiles without sharing your first-party data with third parties.
- Privacy-safe architecture that keeps your customer data protected while still enabling the deep identity resolution needed for personalization, fraud prevention, and omnichannel marketing.
If you are evaluating how an identity graph could work for your business (whether to improve customer recognition, unify fragmented data, or power more relevant experiences), we would love to walk you through how we approach it. Feel free to contact us and start the conversation.