Throttling is the operational lever that keeps a FHIR API responsive when demand outstrips supply. The policy that governs throttling matters more than the mechanism. A blunt policy that caps every client identically produces its own fairness problem; a nuanced policy that carves out capacity per client class keeps interactive traffic honest even when ingestion partners burst.
Naming the policy explicitly is the discipline. For related walkthroughs, more on FHIR for healthcare teams collects the surrounding material.
Why Throttling Beats Auto-Scaling Alone
Auto-scaling adds capacity when demand grows. That works well when the growth is gradual; it fails when the growth is a burst. New instances take seconds to warm up, and by then the interactive workload has already degraded.
Throttling accepts the burst without letting it break the workload. Auto-scaling and throttling are complementary; throttling is what protects the deployment during the seconds auto-scaling needs. A pass through the site's concurrency p99 predictor surfaces the concurrency shape that throttling policy has to defend against.
The Categories the Policy Should Name
A throttling policy that survives production usually names five categories:
- Interactive clinical clients: highest priority, generous headroom.
- Ingestion partners: capped at a defined RPS with burst allowance.
- Analytics extracts: capped at a fraction of total capacity.
- Bulk Data exports: rate-limited to a fraction of I/O.
- Background jobs: lowest priority, generous rate limit but starvation-safe.
Each category gets its own rate limit and priority. Interactive traffic never waits behind ingestion, and ingestion never waits behind analytics.
Rate Limit by Client, Not by Endpoint
Rate limits per endpoint feel intuitive and produce unfair outcomes. A client that fires a chatty pattern hits its endpoint limits faster than a batched client for the same workload.
Rate limits per client capture the actual load the client contributes. Combined with per-endpoint circuit breakers on expensive operations, the two work together to keep the API honest. For the batch case specifically, how batch endpoints skew your concurrency estimate covers the skew that per-endpoint limits miss.
Burst Allowance Matters
Fixed rate limits without a burst window feel harsh to legitimate clients whose workflow legitimately spikes. Token bucket algorithms with a burst allowance let a client burst up to a defined multiple of its steady-state rate and then throttle back.
Every rate-limit implementation should carry both a steady state and a burst. The burst catches the legitimate spikes; the steady state protects the pool. For the pool-side view, background jobs vs interactive requests in the same FHIR pool covers the pool split that throttling should preserve.
The 429 Response Should Be Actionable
Throttled clients should receive a 429 response with a Retry-After header. That header lets the client back off intelligently instead of retrying blindly. Deployments that omit the header force clients to invent their own backoff, and every client invents a different pattern.
The response should also carry a machine-readable reason: rate-limit-exceeded, pool-exhausted, quota-exceeded. Clients that can distinguish among these adjust behavior appropriately.
The Database-Layer Consideration
Throttling at the API preserves the database even when auto-scaling has not caught up. For the deeper database-side story, concurrency limits that come from the database, not the API covers what the throttling actually protects.
Publishing the Policy
Throttling policies that live in tribal knowledge get bypassed. Policies that are published in the developer portal get respected. Every deployment that supports external clients should publish its throttling policy explicitly.
Throttling done well is invisible to legitimate clients and decisive against abusive ones. Every FHIR deployment benefits from the policy stated explicitly rather than emerged accidentally.

Sources
- HL7 FHIR core specification of HTTP interactions defining - HL7 FHIR core specification of HTTP interactions defining 429 handling and rate-limit semantics