Studybox Research

Studybox Research FDA CDRH · Final guidance · 2007

FDA Statistical Guidance for Diagnostic Tests

FDA's rules for how to analyze and report performance of qualitative diagnostic tests, including when sensitivity and specificity may be claimed and why discrepant resolution is discouraged.

Explained

FDA CDRH · Final guidance · 2007

01 What the document says

This guidance covers diagnostic tests with a binary result (positive or negative), whether or not the underlying measurement is quantitative, and applies to data submitted in 510(k)s and PMAs. It was written for both statisticians and non-statisticians and remains the reference FDA reviewers use when checking how IVD clinical performance is calculated and presented.

Its first distinction is between a reference standard and a comparator that is not one. If the new test is compared to an accepted reference standard, results may be reported as sensitivity and specificity with confidence intervals. If the comparator is itself imperfect, for example a cleared predicate or another non-reference method, the guidance asks sponsors to report positive percent agreement and negative percent agreement instead, and not to label those values sensitivity and specificity. It also explains how to present the full two-by-two table, how to handle invalid or indeterminate results, and why the study population should resemble the intended-use population across disease spectrum, demographics, and specimen type.

The guidance pays particular attention to discrepant resolution, the practice of re-testing only the discordant results with a third method and then revising the original counts. FDA explains that this produces biased estimates and asks sponsors not to use it to alter the primary analysis. Where a third method is used, the original agreement table should remain the primary result and any further testing should be reported separately.

02 What it means when you plan a study

  • Decide before enrollment whether your comparator is a reference standard; if it is a cleared predicate, your labeling will carry percent agreement, not sensitivity and specificity, and your sample-size justification should be built on agreement estimates and their confidence intervals.
  • Confidence intervals, not point estimates, drive sample size: the number of positive specimens needed to keep the lower bound above a target is usually the binding constraint and depends heavily on prevalence at the sites you choose.
  • Prospective enrollment of the intended-use population with pre-specified handling of invalid results avoids spectrum bias and the need to defend exclusions later.
  • Plan what happens to discordant results in the protocol: a third method can be run for information, but the primary two-by-two table stays as collected, so the protocol should not promise an adjusted analysis.
  • Collect enough demographic and clinical detail at each site to present performance by subgroup and specimen type, since reviewers expect these breakdowns for IVDs.

03 Pathways it applies to

Source document: Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests (fda.gov).

Where this shows up

Assay Studies Shaped by This Guidance.

Let's talk IVD research

Wherever you fit in the research process, Studybox is here to support you.

Sponsor or clinic: tell us about your study or your site and we'll follow up within one business day.