The concurrency plan sitting in the spreadsheet is not the same as a plan validated by a test. Every deployment eventually finds the gap; the ones that find it during a test are cheaper than the ones that find it during go-live. Naming the scenarios worth running before production is the discipline that surfaces the gap on your calendar rather than the on-call schedule.
Four scenarios cover most of what matters. For related walkthroughs, FHIR background reading collects the surrounding material.
Scenario One: Steady-State Baseline
The baseline test runs the expected workload at the expected concurrency for an extended window. Latency, queue depth, and utilization should sit comfortably within budget. If they do not, the plan needs revision before further testing.
Run the baseline for at least an hour. Short runs miss the slow-burn failures that show up after cache warm-up completes. A pass through the site's concurrency p99 predictor can bracket the expected latency at the baseline concurrency.
Scenario Two: Spike
The spike test doubles or triples the arrival rate for a short window. Queue depth should grow; latency should track behind; utilization should saturate. The metric that matters is recovery time: how long after the spike ends does the system return to baseline behavior.
Deployments that recover in under a minute are usually well-sized. Deployments that take five to ten minutes are marginal; deployments that take longer than that need work. For the queue-side view, queue depth as the leading indicator of trouble covers the recovery signal.
Scenario Three: Sustained Overload
The sustained overload test runs the arrival rate at one hundred twenty percent of the deployment's ceiling for an extended window. The system should degrade gracefully: latency grows, error rate stays under a threshold, downstream systems do not cascade.
This scenario surfaces the ugliest failure modes. Cache stampedes, subscription cascade explosions, and connection pool exhaustion all appear during sustained overload. Fixing them at test time is cheap; fixing them at incident time is not.
Scenario Four: Mixed Workload
Mixed workload testing runs interactive and background traffic together at their expected mix. This is where contention between the two shows up. A bulk export running alongside interactive traffic reveals whether the pool split is holding.
Testing this scenario before go-live is the single change that surfaces the highest-value production failure modes. For the pool-side story, predicting concurrency for a FHIR API before you have real users covers the framing.
Scenario Five (Optional): Batch Endpoint Amplification
Deployments that support batch endpoints should run a scenario that fires realistic-size Bundles at production rates. The persistence tier will see the amplification; if the batches are too large, the tests surface the pain. For the specific skew, how batch endpoints skew your concurrency estimate covers the pattern.
Instrumenting the Tests
Every test should collect the same metrics the production dashboard collects. Latency percentiles, queue depth at every layer, utilization, error rate, and downstream system state. Metrics collected during tests but absent in production produce blind spots; metrics collected in production but absent in tests produce false confidence.
Documenting the Result
The output of each test scenario is a short report: what was tested, what the metrics looked like, what needs fixing before go-live. Test reports are the audit trail future readers use to reconstruct why the deployment was signed off. Every FHIR deployment that runs these scenarios before go-live enters production with numbers that survive contact.

Sources
- Google SRE Workbook chapter on SLOs and load-testing - Google SRE Workbook chapter on SLOs and load-testing methodology applicable to FHIR concurrency