
FHIR analytics at scale runs on Bulk Data Access IG $export feeding a downstream warehouse. The pipeline is straightforward on paper; the operational reality has three chokepoints most teams underestimate.
Chokepoint 1: export scheduling and freshness. Nightly full exports of 10M+ resources take 3-5 hours on typical Postgres-backed FHIR servers. Analytics teams that need same-day freshness need incremental _since exports. HAPI JPA 7.x, Aidbox, and Medplum all support _since cleanly.
Chokepoint 2: NDJSON to warehouse ingestion. BigQuery caps individual files at 4GB, Snowflake's COPY INTO prefers <250MB. Servers that emit unbounded NDJSON files force splitter logic downstream. Use the server's chunking knob if it exists.
Chokepoint 3: terminology expansion at query time. Analytics queries need SNOMED-to-LOINC crosswalks, ICD-10 to SNOMED mappings, etc. Running `$expand` at query time is expensive. Snapshot terminology, join in-warehouse.
Pipeline layout, mid-2026
| Stage | Tool | Cadence |
|---|---|---|
| FHIR bulk export | $export w/ _since |
Nightly + hourly incremental |
| NDJSON ingest | Spark, dbt, BigQuery Load | Nightly |
| Terminology snapshot | $expand batch dump |
Weekly |
| Feature store | dbt or Feast | Nightly rebuild |
| Reporting | Tableau, Looker | Real-time query |
Teams that ship analytics pipelines cleanly have the same architecture: $export on schedule, ingest into warehouse, snapshot terminology, then reporting queries against warehouse (not against FHIR server). Trying to query FHIR directly at reporting scale doesn't work.