FHIR concurrency plans usually target the API layer. Threads per instance, workers per pod, HTTP timeout budgets, requests-per-second ceilings. The layer that actually bounds a mature deployment is the database, not the API, and plans that treat the database as an unbounded downstream discover the truth during an incident.
Naming the database as the constraint is the shift that turns a hopeful plan into a defensible one. For related walkthroughs, the US healthcare interoperability hub collects the surrounding material.
Where the Real Ceiling Lives
Every FHIR API sits in front of a persistence engine. Under load the API's queue depth reflects the persistence engine's throughput, not the API's own capacity. If the database can process one hundred fifty concurrent queries and the API accepts three hundred, the extra one hundred fifty sit in the pool waiting for their database turn.
The visible metric is API queue depth; the invisible one is database connection saturation. A pass through the site's concurrency p99 predictor can project utilization against the underlying database ceiling directly.
Connection Pooling Is Where the Number Sits
The database connection pool is the operational choke point. PgBouncer, pgcat, and the vendor pool inside the persistence engine each cap concurrent database connections at a specific number. Above that number requests queue on the pooler rather than reaching the database.
Sizing plans should name the pool size explicitly and monitor its utilization as a first-class metric. Deployments that leave the pool implicit inherit surprises during traffic peaks.
Long-Running Queries Poison the Pool
A single long-running query holds a pool slot for its full duration. If typical FHIR queries take fifty milliseconds and one bad query takes five seconds, the bad query occupies a slot for one hundred typical queries' worth of time.
Long-running queries in FHIR usually come from search parameter combinations that fall through indexes, wide date-range searches, or Bundle expansions that touch many resources. For the queue-side view of the same failure, queue depth as the leading indicator of trouble covers the shape.
Locks Compound the Effect
Write-heavy workloads add locking to the mix. Transaction Bundle writes hold locks across every entry in the Bundle. A single fifty-entry transaction can serialize behind other writes for its full duration.
Deployments that handle transactional traffic should track lock wait time alongside connection pool utilization. Locks that grow without bound are the shape of contention becoming saturation.
Read Replicas Move the Ceiling
Read-heavy workloads can raise the ceiling by adding read replicas. Every replica adds its own connection pool and its own throughput budget. The plan should name how many replicas serve the workload and how traffic gets routed between them.
Read replicas do not help writes. Write-heavy workloads need a different strategy, usually partitioning or sharding rather than replication.
Throttling at the API Preserves the Database
The mitigation that actually protects the database is throttling at the API. Rate-limiting incoming requests to a level the database can sustain preserves latency for the requests that are accepted. For the specific throttling framing, throttling policies that keep FHIR fair under contention covers the patterns that work.
Throttling feels like it hurts the sender. It hurts less than the alternative, which is every request slowing down when the database saturates.
The Sizing Implication
The sizing plan should carry two numbers: API-layer capacity and database-layer capacity. The smaller number is the deployment's real ceiling. For the wider prediction framing, predicting concurrency for a FHIR API before you have real users covers how to size against both.
Naming the database as the real bound is the small shift that keeps concurrency plans honest.

Sources
- PostgreSQL canonical runtime resource configuration docs - PostgreSQL canonical runtime resource configuration docs applicable to FHIR concurrency limits