Studybox Research Guide
Flex Studies for CLIA Waiver: What FDA Expects and How to Design Them
A flex study deliberately stresses a test system, through wrong volumes, early and late reads, temperature and humidity extremes and the like, to show that results stay correct, or that the device catches the error, under the conditions a waived setting will create.
Field notes
Guide
A flex study for CLIA waiver is a bench study in which the finished device is run under deliberately stressed conditions, such as too little or too much sample, a result read too early or too late, extreme temperature and humidity, a non-level surface or poor lighting. It shows that the test is insensitive to that variation, or that a fail-safe or failure alert catches the error before a wrong result is reported. FDA’s CLIA waiver guidance treats flex studies as half of the evidence that a test has “an insignificant risk of an erroneous result”. Risk analysis is the other half. The stress conditions you test should come straight out of that risk analysis, not from a generic list.
The primary source is FDA’s guidance Recommendations for Clinical Laboratory Improvement Amendments of 1988 (CLIA) Waiver Applications for Manufacturers of In Vitro Diagnostic Devices (issued February 26, 2020; called “the CLIA waiver guidance” below). Its companion, Recommendations for Dual 510(k) and CLIA Waiver by Application Studies (also February 26, 2020), lists flex studies as a required part of a dual submission and refers back to Section IV of the CLIA waiver guidance. Both are nonbinding guidance.
What exactly is a flex study under FDA’s CLIA waiver guidance?
FDA defines flex studies as “studies performed using the device under conditions of operational stress”, performed “to identify potential sources of error as part of the risk assessment” (Appendix B). In the body of the guidance FDA says more about what they are for:
“You should conduct flex studies: studies that stress the operational limits of your test system. Flex studies should be used to validate the insensitivity of the test system to variation under stress conditions. Where appropriate, flex studies should also be used to verify and/or validate the effectiveness of control measures at operational limits.”
The guidance asks applicants to show two things: “(1) the test system design is robust, i.e., insensitive to environmental and usage variation, and (2) that all known sources of error are effectively controlled.” It then splits the work: “flex studies should be used to demonstrate robust design while risk management should be used to demonstrate identification and effective control of error sources, although the two are not mutually exclusive.”
FDA calls this a two-tiered approach:
| Tier | What it contains | What the application should include |
|---|---|---|
| Tier 1: Risk analysis and flex studies | A “systematic and comprehensive risk analysis” of all potential sources of error, plus flex studies that stress the system at and beyond its operational limits | The risk analysis results, a summary of flex study design and results, and the conclusions drawn from them |
| Tier 2: Fail-safe and failure alert mechanisms | The control measures (lock-outs, physical guides, internal and external controls) that reduce each identified risk | A table, for each risk, of the control measure, objective evidence it is implemented, and evidence it works, including under stress |
So flex studies do two jobs: find where results stop being correct, and test whether the controls catch those failures.
Why does a flex study start with hazard and risk analysis?
FDA asks for flex studies “based on the results of the risk analysis and identification of potential problems with sensitivity to environmental or usage variation.” In practice your flex protocol should trace back, condition by condition, to entries in your hazard analysis or FMEA. FDA says it recognizes the ISO 14971 risk-management process and that the guidance “uses the same risk management terminology”. For examples of error sources it also points to CLSI EP18, Risk Management Techniques to Identify and Control Laboratory Error Sources.
The guidance lists potential sources of error to consider, grouped as follows (paraphrased; the full list is in Section IV.A):
- Operator error / human factors: wrong specimen type or volume; incorrect handling, placement, order or amount of reagents; non-level surface; incorrect timing of application, running or reading; misreading, including due to color blindness.
- Specimen, reagent and calibration integrity: collection and handling errors, clotting, interfering substances, bubbles; improperly stored, outdated, mixed or contaminated reagents; calibration stability, including after power failures.
- Hardware, software and electronics: power failure and fluctuation, incorrect voltage, repeated plugging and unplugging, component failure, physical trauma.
- Environmental factors: “temperature, humidity, barometric pressure changes, altitude (if applicable), sunlight, surface angle, device movement, etc.” on reagents, specimens and results, and electrical or electromagnetic interference on instruments.
Two further points matter. First, “you should consider multiple skill levels of users”: the analysis is about waived operators, not R&D staff. Second, FDA states plainly: “We do not recommend training as a sole means of mitigating potential sources of harm.” If a flex study finds a failure mode, the expected fix is a design change, a fail-safe, or a validated failure alert. Adding a sentence to the instructions is not enough on its own.
Which stress conditions should a flex study cover?
Cover every condition your risk analysis flags as plausible in a waived setting: for most unitized and reader-based tests, sample and buffer volume, read time, temperature and humidity, storage, surface level, disturbance and lighting, plus power and electromagnetic conditions for instruments.
The CLIA waiver guidance gives two worked examples in its Table 1. In the storage example, the labeled storage is 2–4 °C and kits were stored at 0, 2, 10, 25 and 37 °C, which showed failure when frozen or held at 25 °C for more than 3 days. In the drop-count example, the procedure calls for 3 drops and 1 through 6 drops were tested, which showed erroneous results below 2 or above 5 drops. In both cases the next step is a validation study showing that a fail-safe or failure alert flags exactly those failing conditions.
For specific stress levels, the most detailed public FDA text is the Appendix A of FDA’s pandemic-era EUA template for home-use molecular and antigen COVID-19 tests. It was written for one analyte and pathway, but its designs follow Section IV’s logic, and it points developers to CLIA waiver decision summaries for “alternative flex study designs”. The table below combines the CLIA waiver guidance’s error sources with those designs:
| Stress condition | Where FDA names it | Example design given by FDA |
|---|---|---|
| Sample volume | CLIA waiver guidance (incorrect volume); EUA home template | Test at half and twice the labeled volume, plus the maximum that can be added (e.g., 5, 10, 20 µL for a 10 µL procedure); add intermediate levels if a limit fails |
| Reagent / buffer volume | CLIA waiver guidance Table 1; EUA home template | 1–6 drops for a 3-drop procedure (CLIA guidance); 1, 2, 3, 4 drops and the whole bottle for a 2-drop procedure (EUA template) |
| Read time | CLIA waiver guidance (incorrect timing of reading); EUA home template | From four-fold below to three-fold above the labeled time; for a 20-minute read, at least 5, 10, 15, 20, 30 and 60 minutes |
| Delays between steps | EUA home template | Delay in sample testing and in operational steps |
| Temperature and humidity (operating) | CLIA waiver guidance (environmental factors); EUA home template | 40 °C / 95% RH (hot, humid) and 5 °C / 5% RH (cold, dry) |
| Storage excursions | CLIA waiver guidance Table 1 | Storage across and outside the labeled range (0–37 °C in the example) |
| Swab elution / mixing | EUA home template | From no mixing to vigorous shaking with bubbles, plus intermediate swirling |
| Lighting | CLIA waiver guidance (sunlight; color-blindness reading error); EUA home template | Fluorescent, incandescent and natural light |
| Surface angle / orientation | CLIA waiver guidance (non-level surface, surface angle); EUA home template | Run off-level or in the wrong orientation as the risk analysis indicates |
| Disturbance during the run | CLIA waiver guidance (device movement, physical trauma); EUA home template | Moving or dropping the device mid-run, unplugging, phone call during an app-run test |
| Power and electronics | CLIA waiver guidance (power failure, fluctuation, voltage, plugging) | Power interruption, battery failure, repeated plugging and unplugging |
| Electromagnetic interference | CLIA waiver guidance; EUA home template | Cell phones, Bluetooth, Wi-Fi and other equipment expected in the use environment |
| Altitude / barometric pressure | CLIA waiver guidance (“if applicable”) | Device-specific; for glucose meters FDA asks for testing to at least 10,000 ft with a real pressure change (SMBG OTC guidance) |
Meter-based systems have their own conventions. FDA’s 2020 OTC blood glucose guidance lists flex studies for that class, including “the four extreme temperature and humidity combinations”, short-sample detection, sample perturbation and intermittent sampling. They are glucose-specific, but they show the expected approach: build the flex design around the failure modes of the device in front of you.
How should the flex study itself be designed?
Run each condition with samples that would reveal a shift, include enough replicates to see failures, compare against the nominal condition, and keep going until the device fails or the condition is clearly beyond anything a real operator would do.
The EUA home template gives the clearest public numbers. Flex studies “should be performed in-house by staff who have been trained in the use of the test”, using “a negative sample and a low positive sample (at 1.5 - 2 times LoD) for each condition”, “conducted to the point of failure to determine the maximum deviation that will still generate accurate results”, with “3 replicates per condition per sample concentration”, and line data provided. Treat those as stated expectations for one analyte, not a universal rule. Quantitative tests use levels around the medical decision points and compare against results under nominal conditions.
Why the low positive sits above LoD, not at it. A sample at the LoD is, by definition, detected only most of the time. If a sample’s true hit rate under nominal conditions is 95%, the chance that all three replicates come back positive is 0.95³ = 0.857. In other words, about one run in seven (14.3%) would show at least one negative with no flex effect at all. At a 99% hit rate the chance of a spurious miss in three replicates falls to 3%. Testing at 1.5–2× LoD keeps the nominal hit rate high, so a negative under stress means something. Whatever level you pick, write down before testing what counts as a failure (for example, any discordant replicate triggers an extended set at that condition). Otherwise you will be making judgement calls after you have seen the data.
Why “to the point of failure” matters. Testing only the labeled limits shows the device works inside the window, not how close the cliff is. FDA’s Table 1 examples find the boundary (under 2 and over 5 drops) because that boundary is what the failure-alert validation then has to cover.
Validate the controls under the same stress. Section IV.C of the guidance asks for studies that validate “all fail-safe and failure alert mechanisms … under conditions that stress the device”. FDA also cautions that procedural controls “generally provide limited problem detection and, by themselves, are generally not sufficient to serve as a failure alert mechanism”, and that flex and validation studies “should evaluate the sensitivity of internal control reagents to all applicable test system errors.” If your control line only confirms that liquid flowed, the labeling has to say so (Section VI.C).
How do flex studies relate to near-cutoff and reproducibility studies?
They answer different questions, and one cannot stand in for the other. A flex study asks whether the device stays correct, or flags the error, when a condition is pushed past normal use. A reproducibility study asks how much results vary between untrained operators, sites, days and lots under normal use. The comparison study asks whether untrained operators get the same answers as trained operators or the comparator on real patient samples.
| Flex study | Reproducibility study | Comparison study (near-cutoff samples) | |
|---|---|---|---|
| Question | Is the result robust, or the error caught, under stress? | How much do results vary under normal use? | Do untrained operators match trained operators or the comparator? |
| Who runs it | Trained in-house staff (per the EUA template) | Untrained operators at ≥3 of the comparison sites (dual guidance) | ≥9 untrained operators at ≥3 intended-use sites (CLIA waiver guidance) |
| Samples | Contrived negative and low positive, or levels around decision points | True negative, high negative (near C5), low positive (near C95), moderate positive, for a clinical cutoff (dual guidance) | Prospective patient samples, including some near the cutoff; archived or surrogate samples generally ≤ one third |
| Output | Failure boundaries; evidence that controls detect them | Agreement or precision across sources of variation | Agreement estimates with confidence intervals |
The dual guidance defines the near-cutoff levels: “C5 is a sample concentration which yields a positive result 5% of the time … and C95 is a sample concentration which yields a positive result 95% of the time.” Those samples are built to be uncertain, so a flex study cannot run on them. The flex study needs a sample whose answer is stable at nominal conditions, so that any change can be blamed on the stress.
The two connect through risk. If flex testing shows that late reads turn high negatives into faint positives, expect the same in the reproducibility panel’s near-C5 samples when an operator is distracted. Fix it before that study, not after.
There is one situation where flex studies carry more weight. Under Option 3 of the CLIA waiver guidance, “flex and human factors engineering studies may provide sufficient assurance” in place of a comparison study. FDA says this generally applies only where specimen collection is always done by a professional or always by the patient, other pre-analytical steps are very simple, and intended-use patient populations are sufficiently similar. It also applies to some modifications of previously waived tests. Option 3 is a Pre-Submission conversation, not a default.
What should a flex study protocol include?
Use the checklist below to review a flex protocol before it is executed. Each row maps to text in the CLIA waiver guidance or, where noted, the EUA home template.
| Protocol element | What to check | Basis |
|---|---|---|
| Traceability | Every condition maps to a hazard or failure mode in the risk analysis, and every high-risk failure mode has a condition | CLIA guidance IV.A |
| Device configuration | Final design, final reagent formulation, final QRG | EUA template (“final design/format”) |
| Operators | Trained in-house staff, identified | EUA template |
| Samples | Negative plus low positive (e.g., 1.5–2× LoD), or quantitative levels at decision points; matrix matches intended specimen | EUA template; CLIA guidance V.C(4) on matrix |
| Nominal arm | Each condition compared to the labeled procedure run in parallel | SMBG OTC guidance (comparison to nominal) |
| Levels | Bracket the labeled value and continue to the point of failure | CLIA guidance Table 1; EUA template |
| Replicates | Stated per condition per level (EUA template: 3), with a pre-specified rule for discordant replicates | EUA template |
| Acceptance criteria | Pre-defined: expected qualitative result, or allowable bias against nominal | EUA template (“pre-defined study protocol”) |
| Control response | Internal and external controls observed under each stress; fail-safe/alert behavior recorded | CLIA guidance IV.C |
| Readout | For visual reads, lighting and color-perception conditions included; reader blinded to condition where feasible | CLIA guidance IV.A |
| Data | Line data for every replicate, including invalids | EUA template |
| Outcome handling | Each failure leads to a design control, fail-safe or validated alert; labeling-only fixes justified | CLIA guidance IV.B (“training as a sole means”) |
What are the most common flex study deficiencies?
The gaps are usually structural, not statistical; each reverses a specific recommendation in the guidance.
- Conditions not traced to the risk analysis. A generic list that leaves out the device’s own failure modes (a half-seated cartridge, an ampoule squeezed twice). FDA asks you to “consider any other potential system failures that may be specific to your device.”
- Stopping at the labeled limits. Testing only inside the claimed range shows nothing about margin. FDA’s own examples run past the point where results fail.
- The wrong sample level. Only moderate or strong positives, which hide shifts near the cutoff. Or samples at the LoD, where spurious misses make the results impossible to interpret.
- Read-time window not bracketed on both sides. Early reads produce false negatives and late reads can produce false positives. Both directions need data, and both feed the reading-time instruction in the QRG.
- Failures mitigated by labeling alone. FDA does not recommend training as the sole mitigation, and the guidance prefers fail-safe mechanisms “whenever it is technically practicable”.
- Controls not tested under stress. A control line that still appears with half the required buffer is not detecting that error. The guidance asks for control validation “under conditions that stress the device” and for the control’s limitations to be stated in labeling.
- Visual-read conditions ignored. Lighting and color perception both appear in FDA’s lists, and the guidance’s QRG appendix calls for a warning addressing color blindness “when waived tests use color-coded reagents and/or endpoints.”
- Design or QRG changed after flex testing. Later changes need a documented rationale or repeat testing.
How Studybox approaches this
Flex studies are usually the manufacturer’s bench work, but they shape what the field studies have to show. Studybox Research works only on in vitro diagnostics, and its team has supported 55+ FDA regulatory clearances across 510(k), dual 510(k)/CLIA waiver, OTC and EUA pathways. On waiver programs we read the flex and risk-analysis results alongside the CLIA waiver or dual submission protocol, so that any failure boundary found on the bench is reflected in the QRG, the operator questionnaire and the near-cutoff panel before untrained operators see the device. Examples include a point-of-care PT/INR system tested with untrained operators and a reader-based influenza A/B antigen test.
The field side runs on 100+ pre-qualified U.S. urgent care, physician office and point-of-care sites, with typical activation in about four weeks against a 3–6 month industry average. If you are planning a waiver submission and want a second read of how your flex results connect to the clinical protocol, contact us.
Studybox Research Guide October 5, 2026