Clinia
Concepts

Entity Resolution

How the Clinia Context Engine deduplicates clinical entities across multiple sources using a 4-layer escalating pipeline.

Entity Resolution

A complex patient record typically has the same condition, medication, or allergy represented multiple times across different sources, with different codes, different display text, and sometimes contradictory details.

Entity resolution is the process of deciding: are these two records the same real-world thing?

The 4-Layer Pipeline

The Clinia Context Engine uses an escalating cascade that only moves to more expensive methods when cheaper ones can't decide.

LayerMethodHandles
1Deterministic code matchingEntities sharing standard codes (SNOMED, RxNorm, ICD-10, LOINC)
2NLP normalization + fuzzy matchingSame concept with different display text or minor variation
3Embedding similaritySemantically similar entities with no shared codes
4LLM adjudicationGenuinely ambiguous cases requiring clinical judgment

A pair of entities only escalates to the next layer when the current layer cannot confidently resolve it. Layer 4 fires for at most ~15 pairs per patient to ensure cost-effectiveness. Pairs that Layer 4 cannot confidently resolve persist as two separate facts — there is no additional merge pass after Layer 4.

In practice: ~40% of pairs resolve in Layer 1 for free. ~70% resolve before reaching LLM. The LLM sees only hard cases where clinical judgment is genuinely needed.

Provenance

Every merged entity carries a full audit trail (see Provenance and Auditability for the complete structure):

{
  "sources": [
    { "type": "fhir", "origin": "FHIR Bundle", "reliability": 0.85, "ref": "Condition/abc" },
    {
      "type": "cda",
      "origin": "Summary_20230907.xml",
      "reliability": 0.8,
      "ref": "observation/456"
    }
  ],
  "conflicts": [],
  "resolvedBy": "deterministic-code",
  "reasoning": "Both entities share SNOMED code 44054006.",
  "confidence": 0.9
}

When sources conflict (e.g., different medication doses), the conflict is recorded with both values and the resolution strategy used. Every merge decision is fully auditable.

Conflict Resolution

When sources disagree on an attribute value, the pipeline uses one of these strategies:

StrategyWhen used
CorroborationSelect the value agreed upon by the majority of sources
RecencyPrefer the most recent source
ReliabilityWeight sources by their reliability score
LLM judgmentComplex conflicts where an LLM evaluates both values with full patient context

Type Correction

Before deduplication runs, the pipeline corrects entities that the source system encoded under the wrong type. Conditions coded Z80–Z84 or whose display names a relative ("family history of colon cancer") are moved to the family history subgraph; conditions naming an implant or appliance ("pacemaker in situ") become devices. The correction is applied as the raw fact is staged, so the entity is deduplicated under its corrected type.

Facts vs. memories

The output of entity resolution is a graph of facts: deduplicated, source-merged clinical entities (conditions, medications, allergies, and so on) that represent what is structurally known about a patient. Facts are distinct from memories, which capture unstructured conversational statements and are never merged into this graph.

Auditing Resolution Decisions

The full resolution report (GET /v1/patients/{patientId}/resolution) returns every entity in the graph with its complete provenance trace. Use this to debug unexpected merge or non-merge outcomes. See Audit Entity Resolution.

For a higher-level reconciliation summary (stats, merge groups, inferred relationships), use GET /v1/patients/{patientId}/reconciliation. See Review the Reconciliation Summary.

On this page