FHIR Analytics Pipelines: Bulk Data to Warehouse in Practice

FHIR Analytics Pipelines: Bulk Data to Warehouse in Practice

Diagram: FHIR analytics pipeline — chokepoints from FHIR server through ingest to warehouse and reporting

FHIR analytics at scale runs on Bulk Data Access IG $export feeding a downstream warehouse. The pipeline is straightforward on paper; the operational reality has three chokepoints most teams underestimate.

Chokepoint 1: export scheduling and freshness. Nightly full exports of 10M+ resources take 3-5 hours on typical Postgres-backed FHIR servers. Analytics teams that need same-day freshness need incremental _since exports. HAPI JPA 7.x, Aidbox, and Medplum all support _since cleanly.

Chokepoint 2: NDJSON to warehouse ingestion. BigQuery caps individual files at 4GB, Snowflake's COPY INTO prefers <250MB. Servers that emit unbounded NDJSON files force splitter logic downstream. Use the server's chunking knob if it exists.

Chokepoint 3: terminology expansion at query time. Analytics queries need SNOMED-to-LOINC crosswalks, ICD-10 to SNOMED mappings, etc. Running `$expand` at query time is expensive. Snapshot terminology, join in-warehouse.

Pipeline layout, mid-2026

Stage Tool Cadence
FHIR bulk export $export w/ _since Nightly + hourly incremental
NDJSON ingest Spark, dbt, BigQuery Load Nightly
Terminology snapshot $expand batch dump Weekly
Feature store dbt or Feast Nightly rebuild
Reporting Tableau, Looker Real-time query

Teams that ship analytics pipelines cleanly have the same architecture: $export on schedule, ingest into warehouse, snapshot terminology, then reporting queries against warehouse (not against FHIR server). Trying to query FHIR directly at reporting scale doesn't work.