Studybox Research Guide
CLIA Waiver Study Sample Size: Working Back From the Wilson Score Lower Bound
For a qualitative CLIA waiver study, the binding constraint is usually the number of comparator-positive specimens needed for the lower bound of the two-sided 95% confidence interval to clear the acceptance criterion. Prevalence and the evaluable rate then turn that number of positives into a total enrollment target.
Field notes
Guide
For a binary qualitative test, the size of a CLIA waiver or dual 510(k)/CLIA waiver comparison study is driven less by total subjects than by the number of comparator-positive specimens. FDA asks for agreement or accuracy estimates with two-sided 95% confidence intervals, and the acceptance decision usually rests on the lower bound of that interval for positive percent agreement (or sensitivity). Work out how many positives you need for the lower bound to clear your criterion at the performance you realistically expect, then divide by prevalence and by the fraction of enrolled subjects you expect to be evaluable. That gives the enrollment target.
Below: where those inputs come from in FDA’s documents, the Wilson score arithmetic, and worked tables for a first-pass plan before a Pre-Submission.
What do FDA’s documents actually require for sample size?
FDA’s CLIA waiver guidance does not set a single sample size or a universal acceptance criterion. The 2020 guidance, Recommendations for CLIA Waiver Applications for Manufacturers of In Vitro Diagnostic Devices, says plainly that “the appropriate acceptance criteria for the studies performed using the design options described above will vary from test to test,” depending on intended use, context of use and the probable benefits and risks of waived use. It adds that the minimum agreement between untrained and trained users “should generally be higher for a test for which erroneous results in waived settings are associated with a higher extent of probable patient risk/harm.”
What the guidance does fix are structural minimums that put a floor under the study:
| Element | Recommendation | Source |
|---|---|---|
| Testing sites | ”a minimum of three sites that are representative of both the intended use patient population and the intended operators in CLIA-waived settings” | CLIA waiver guidance, Section V.C(1) |
| Untrained operators | ”1-3 untrained operators at each site and at least nine (9) untrained operators across all sites” | CLIA waiver guidance, Section V.C(2) |
| Enrollment pattern | Prospective specimens “collected from consecutive patients over one month,” which may be limited to two weeks depending on site and prevalence | CLIA waiver guidance, Section V.C(4) |
| Archived or surrogate specimens | ”In general … should not comprise greater than one third of the total study samples” | CLIA waiver guidance, Section V.C(4) |
| Per-operator specimens (dual submission, binary qualitative) | “Each untrained operator should run the candidate test with a minimum of 5 samples that are positive by the CM and 5 samples that are negative by the CM” | Dual 510(k) and CLIA Waiver guidance, Section V.A(2) |
| Invalid results | Report initial and final invalid percentages “with a 95% two-sided confidence interval and then exclude invalid results from calculations of the test performance characteristics” | CLIA waiver guidance, Section V.C(6) |
Combining two of those recommendations gives a simple arithmetic floor for a dual submission: nine untrained operators each testing at least five comparator-positive specimens is at least 45 positives. That floor is rarely what determines the final number. The confidence interval does.
Where FDA has written numeric performance criteria, they appear in device-specific documents rather than in the CLIA waiver guidances. One example is the special control for influenza antigen tests at 21 CFR 866.3328: against an FDA-cleared nucleic acid test or other accepted comparator other than culture, positive percent agreement “must be at the point estimate of at least 80 percent with a lower bound of the 95 percent confidence interval that is greater than or equal to 70 percent,” and negative percent agreement at a point estimate of at least 95% with a lower bound of at least 90%. Criteria like these, a point estimate plus a lower bound, are the shape most qualitative acceptance criteria take, which is why the lower bound is the number to plan around.
Why is the lower bound, not the point estimate, what sizes the study?
Because the point estimate does not depend on sample size and the lower bound does. Ninety-five agreements out of 100 and 19 out of 20 are both 95.0%, but the second gives far less assurance, and the only lever on that assurance is the denominator.
For PPA or sensitivity, the denominator is comparator-positive (or reference-positive) specimens. For NPA or specificity, it is comparator-negative specimens. In most intended-use populations negatives are plentiful, so NPA’s lower bound tightens quickly and the positive arm sets the enrollment.
Which confidence interval method does FDA use?
FDA’s Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests (2007) recommends reporting sensitivity and specificity, or positive and negative percent agreement, “and their two-sided 95 percent confidence intervals,” and its worked examples use “two-sided 95% score confidence intervals,” referring readers to Altman et al. (2000) and CLSI EP12 for the calculation, with exact (Clopper-Pearson) intervals mentioned as an alternative. The score interval for a single proportion is the Wilson score interval.
The guidance’s own examples reproduce exactly with the Wilson formula. We checked them in Python:
| FDA 2007 example | Reported 95% score CI | Wilson score CI (computed) |
|---|---|---|
| Sensitivity 44/51 = 86.3% | 74.3% to 93.2% | 74.3% to 93.2% |
| Specificity 168/169 = 99.4% | 96.7% to 99.9% | 96.7% to 99.9% |
| PPA 40/44 = 90.9% | 78.8% to 96.4% | 78.8% to 96.4% |
| NPA 171/176 = 97.2% | 93.5% to 98.8% | 93.5% to 98.8% |
For x agreeing results out of n, with p̂ = x/n and z = 1.96, the Wilson interval is:
(p̂ + z²/2n ± z·√[p̂(1 − p̂)/n + z²/4n²]) / (1 + z²/n)
Unlike the simple normal-approximation interval, it behaves sensibly at the high agreement rates typical of IVD studies and still gives a lower bound below 100% when every result agrees. The Clopper-Pearson exact interval is more conservative and will give somewhat lower bounds at the same counts, so state the method in the protocol and use the same one at analysis.
How many positives do you need for a given lower bound?
It depends on how many positives the test misses. The table below gives the Wilson lower bound for PPA (or sensitivity) by the number of comparator-positive specimens and the number missed. All values were computed with the formula above (z = 1.959964).
| Positives (n) | 0 missed | 1 missed | 2 missed | 3 missed | 5 missed | 10 missed |
|---|---|---|---|---|---|---|
| 30 | 88.6% | 83.3% | 78.7% | 74.4% | 66.4% | 48.8% |
| 50 | 92.9% | 89.5% | 86.5% | 83.8% | 78.6% | 67.0% |
| 75 | 95.1% | 92.8% | 90.8% | 88.9% | 85.3% | 77.2% |
| 100 | 96.3% | 94.6% | 93.0% | 91.5% | 88.8% | 82.6% |
| 150 | 97.5% | 96.3% | 95.3% | 94.3% | 92.4% | 88.2% |
| 200 | 98.1% | 97.2% | 96.4% | 95.7% | 94.3% | 91.0% |
| 300 | 98.7% | 98.1% | 97.6% | 97.1% | 96.2% | 94.0% |
The same arithmetic, viewed by observed agreement rate rather than misses:
| Positives (n) | Observed 95% | Observed 96% | Observed 98% |
|---|---|---|---|
| 50 | n/a (not a whole count) | 48/50, lower bound 86.5% | 49/50, lower bound 89.5% |
| 100 | 95/100, lower bound 88.8% | 96/100, lower bound 90.2% | 98/100, lower bound 93.0% |
| 200 | 190/200, lower bound 91.0% | 192/200, lower bound 92.3% | 196/200, lower bound 95.0% |
| 300 | 285/300, lower bound 91.9% | 288/300, lower bound 93.1% | 294/300, lower bound 95.7% |
| 500 | 475/500, lower bound 92.7% | 480/500, lower bound 93.9% | 490/500, lower bound 96.4% |
Two things stand out. First, the gap between the point estimate and the lower bound shrinks slowly: going from 100 to 500 positives at 95% observed agreement raises the lower bound by only about four points. Second, at small n a single additional miss moves the lower bound by three to five points, which matters when a study has 30 to 50 positives.
How do you work backward from a target lower bound?
Pick the acceptance criterion, pick a realistic expected performance, and find the smallest n that clears the criterion. There are two ways to do this, and they give very different answers.
The deterministic approach assumes the study will observe exactly the expected performance. In the table below, the observed count is the expected rate times n, rounded down, and n is the smallest number of positives for which the Wilson lower bound reaches the target.
| Target lower bound | Expected PPA | Smallest n (observed) |
|---|---|---|
| 70% | 85% | 39 (33/39) |
| 70% | 90% | 25 (22/25) |
| 70% | 95% | 15 (14/15) |
| 80% | 90% | 68 (61/68) |
| 80% | 95% | 33 (31/33) |
| 90% | 95% | 140 (133/140) |
| 90% | 98% | 69 (67/69) |
The probabilistic approach accepts that the observed count is random. If the true PPA is 90%, a study of 68 positives will observe fewer than 61 agreements a substantial fraction of the time. Treating the number of agreements as binomial, we computed the exact probability that the Wilson lower bound clears the target (the study’s power) and found the smallest n giving 80% and 90% power.
| Target lower bound | True PPA | First n with 80% power | 80% power from here on | First n with 90% power | 90% power from here on |
|---|---|---|---|---|---|
| 70% | 85% | 60 | 69 | 81 | 89 |
| 70% | 90% | 30 | 35 | 39 | 43 |
| 80% | 90% | 100 | 112 | 130 | 142 |
| 80% | 95% | 40 | 40 | 47 | 54 |
| 90% | 95% | 217 | 254 | 291 | 315 |
| 90% | 98% | 69 | 84 | 84 | 99 |
With discrete counts, power does not rise smoothly with n; it saw-tooths, so a slightly larger n can occasionally have slightly lower power. The “from here on” columns give the n beyond which power stayed at or above the stated level for every larger n we checked (the next 50). If you want a size that is robust to losing a few specimens, plan to that number.
Planning on the deterministic number means accepting a large chance of failure. For a 90% true PPA against an 80% lower-bound criterion, the deterministic answer is 68 positives, but a study of 68 positives clears the criterion only about 63% of the time. Reaching 80% power takes 100 positives, and 112 to stay above 80% for every larger n.
These calculations assume the expected performance is right. It is the most important and least certain input. Base it on analytical data, earlier clinical data with trained operators, and the comparator’s own performance, and discuss it with FDA in a Pre-Submission, which both guidances recommend.
How does prevalence turn positives into total enrollment?
Divide the positives you need by the expected positivity rate in the enrolled population, then by the fraction of enrolled subjects you expect to be evaluable:
Total enrolled = positives needed / (prevalence × evaluable fraction)
| Positives needed | Prevalence | 95% evaluable | 90% evaluable |
|---|---|---|---|
| 50 | 30% | 176 | 186 |
| 50 | 10% | 527 | 556 |
| 50 | 5% | 1,053 | 1,112 |
| 100 | 30% | 351 | 371 |
| 100 | 20% | 527 | 556 |
| 100 | 10% | 1,053 | 1,112 |
| 100 | 5% | 2,106 | 2,223 |
| 100 | 3% | 3,509 | 3,704 |
Prevalence dominates. Halving it doubles enrollment, while the difference between 90% and 95% evaluability is about 5%. Three practical points follow.
Use positivity in the population you will actually enroll, by season. For respiratory targets, positivity among symptomatic patients presenting at urgent care in peak season can be several times the off-season rate. The CLIA waiver guidance’s recommendation of consecutive enrollment over a month, or two weeks where prevalence and site circumstances justify it, means the calendar window is a design input, not an afterthought.
Negatives come for free, up to a point. At 10% prevalence, enrolling 1,112 subjects for 100 positives also yields roughly 900 evaluable negatives, far more than an NPA lower bound typically needs. That surplus is the cost of a prospective, all-comers design, and it is what gives the study its representativeness.
Enrichment has limits. When positives are rare, the CLIA waiver guidance allows archived or surrogate specimens to supplement prospective ones, but “in general” no more than one third of total study samples, appropriately justified and ideally discussed in a Pre-Submission. FDA’s 2007 Statistical Guidance also notes that enriched positives may be “inappropriate for pooling” with prospective positives and recommends consulting FDA.
How should dropouts, invalids and protocol deviations be budgeted?
As a separate loss term, estimated per stage. Subjects are lost between enrollment and analysis for several distinct reasons, and each has its own rate:
- Eligibility failures discovered after consent.
- Pre-analytical losses: comparator specimens that are lost, leak, exceed stability or arrive outside temperature limits.
- Protocol deviations that render a result non-evaluable under the pre-specified analysis plan.
- Invalid results on the candidate test after any permitted retest.
The CLIA waiver guidance treats invalids explicitly: report the number of initial invalids, retests and final invalids for each operator, calculate the initial and final invalid percentages with 95% two-sided confidence intervals, exclude invalids from the performance calculations, and give “a rationale as to why the observed percentage of invalid results is clinically acceptable.” Because invalids are excluded from the denominator, they reduce evaluable positives directly, and a high invalid rate is itself a finding reviewers will examine.
FDA’s 2007 Statistical Guidance adds an accounting requirement: report the number of subjects planned, tested, used in the final analysis and omitted. Equivocal or indeterminate results should not simply be dropped. The guidance suggests one option is to report performance twice, once counting equivocals as positive and once as negative, and recommends consulting FDA statisticians.
The published case studies on this site show how much the evaluable fraction varies. In one point-of-care molecular RSV study, roughly 93% of enrolled subjects were evaluable after protocol-deviation rejections and invalids. In a larger point-of-care STI study, about 94% of enrolled subjects entered the performance evaluation, and the initial invalid rate was around 7%. Budget the loss rate from your own pilot or analytical data, not from a best case.
How should discordant results be handled in the analysis?
Report the original two-by-two table and calculate agreement from it, without revising results based on discordant testing. FDA’s 2007 Statistical Guidance states that you “should not use outcomes that are altered or updated by discrepant resolution to estimate the sensitivity and specificity of a new test or agreement between a new test and a non-reference standard.” It explains that apparent agreement calculated from a revised table can only rise, since results move from disagreement cells to agreement cells but never the reverse, and that “using a coin flip as the resolver will also improve apparent agreement.”
Discordant testing with a second method is still common and can be informative as a descriptive footnote, for example noting how many false negatives were also negative on an alternative molecular assay. Several of the case studies on this site report exactly that kind of footnote alongside an unrevised primary analysis. What it cannot do is change the counts that the lower bound is computed from. For sample-size purposes, this means planning on the unrevised table. Do not plan on discordant testing “recovering” misses.
What else changes the number?
Per-site and per-operator reporting. The CLIA waiver guidance asks for results “by each intended site and if appropriate, overall,” and the dual guidance’s minimum of five comparator-positive specimens per untrained operator means positives must be spread across operators, not concentrated at the highest-prevalence site. A study can reach its total positive count and still fall short at the operator level.
Multiple targets. A multiplex panel needs enough positives for each claimed analyte. The rarest target sets enrollment, and the common targets accrue a surplus.
Specimen types and populations. Separate claims, such as distinct specimen types or symptomatic and asymptomatic populations, may each need adequately sized estimates.
Near-cutoff specimens. The CLIA waiver guidance asks that qualitative studies include specimens near the cutoff. These are the specimens most likely to produce untrained- operator errors, so they belong in the plan rather than appearing by chance.
A worked planning example
Suppose a point-of-care antigen test for a respiratory target is expected to have a true PPA of 90% against an FDA-cleared molecular comparator, and the sponsor plans to an acceptance criterion of a 95% lower bound of at least 80% (a criterion to be confirmed with FDA, used here only as an illustration).
- Positives for 80% power: 100 comparator positives (from the power table above; 112 if you want the more robust figure).
- Operator floor: 9 operators × 5 positives = 45, comfortably below 100.
- Prevalence: expected positivity of 20% in symptomatic patients during the enrollment window.
- Evaluable fraction: 90%, after eligibility, specimen and invalid losses.
- Enrollment: 100 / (0.20 × 0.90) = 556 subjects.
If the season underdelivers and positivity is 10%, the same study needs 1,112 subjects. That sensitivity of enrollment to prevalence is the main reason to enroll across several geographies and to decide in advance how the study will respond to a slow season, for example by adding sites rather than extending the window.
How Studybox approaches this
We size waiver and dual-submission studies from the positive arm backward: agree on the acceptance criterion and expected performance with the sponsor, compute the positives needed with the interval method named in the protocol, and set enrollment from realistic, seasonal positivity and an explicit evaluability budget. The goal is a number that a reviewer can trace and that survives a slower-than-hoped season.
Our network of 100+ pre-qualified U.S. sites spans urgent care, physician office and point-of-care settings, which is where untrained intended-use operators and symptomatic patients are. Prospective collection runs under a pre-approved IRB protocol and can start in as little as one week, and typical study activation is about four weeks, so adding sites to recover positives is a practical option rather than a re-plan. More on our CLIA waiver studies and dual 510(k)/CLIA waiver studies, or contact us to discuss a specific design.
Studybox Research Guide October 5, 2026