Studybox Research

Studybox Research Guide

PPA and NPA vs Sensitivity and Specificity: Which One FDA Expects

PPA and NPA are calculated exactly like sensitivity and specificity. The difference is the comparator: against a reference standard you report sensitivity and specificity, and against anything less, FDA expects positive and negative percent agreement.

← All guides October 5, 2026 11 min read

Field notes

Guide

Positive percent agreement (PPA) and negative percent agreement (NPA) use the same arithmetic as sensitivity and specificity: the proportion of comparator-positive specimens the new test calls positive, and the proportion of comparator-negative specimens it calls negative. What differs is what the comparator is. FDA’s 2007 Statistical Guidance reserves “sensitivity” and “specificity” for comparisons against a reference standard, the best available method for establishing whether the target condition is present. When the comparator is a non-reference standard, such as a predicate device or another cleared assay, FDA recommends the terms PPA and NPA, because the numbers measure agreement, not correctness.

The choice is not cosmetic. It determines which claims can appear in labeling, whether predictive values can be reported at all, and which acceptance criteria apply.

How does FDA define PPA, NPA, sensitivity and specificity?

FDA’s Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests (March 2007) defines all four from a two-by-two table of new-test results against the comparator.

Against a reference standard, with TP, FP, FN and TN as true and false positives and negatives:

  • Sensitivity = 100% × TP / (TP + FN)
  • Specificity = 100% × TN / (FP + TN)

Against a non-reference standard, where the cells are labeled a (both positive), b (new test positive, comparator negative), c (new test negative, comparator positive) and d (both negative):

  • Positive percent agreement = 100% × a / (a + c)
  • Negative percent agreement = 100% × d / (b + d)
  • Overall percent agreement = 100% × (a + d) / (a + b + c + d)

The formulas are identical. The guidance explains the distinction: when a non-reference standard is used, “the same numerical calculations are made, but the estimates are called positive percent agreement and negative percent agreement, rather than sensitivity and specificity. This reflects that the estimates are not of accuracy but of agreement of the new test with the non-reference standard.”

Sensitivity / specificityPPA / NPA
ComparatorReference standard (single method or a defined combination)Non-reference standard (predicate, cleared assay, other method)
What it estimatesAccuracy: how often the new test is correctAgreement: how often the new test matches the comparator
Predictive values, likelihood ratiosCan be calculated (PPV and NPV also depend on prevalence)“Cannot be computed,” per FDA 2007
Direction mattersDefined against condition statusAgreement of new test with comparator differs from the reverse; state the calculation

What counts as a reference standard?

FDA uses the STARD definition: a reference standard is “considered to be the best available method for establishing the presence or absence of the target condition.” It can be a single test or a combination of methods, including clinical follow-up. When it is a combination, “the algorithm specifying how the different results are combined to make a final positive/negative classification … is part of the standard.”

The guidance is candid that this is a judgment, “established by opinion and practice within the medical, laboratory, and regulatory community,” and that sometimes no consensus reference standard exists or a standard is known to be wrong for part of the population. In those situations it recommends consulting FDA on the choice before the study begins.

When does FDA expect agreement language instead of accuracy language?

FDA’s 2007 guidance sets out four cases, in order of preference:

  1. A reference standard is available: use it to estimate sensitivity and specificity.
  2. A reference standard is available but impractical: use it on a subset of subjects and calculate sensitivity and specificity adjusted for verification bias, ideally with a CDRH statistician involved. The guidance notes that naive calculations in this design “would give biased estimates.”
  3. No reference standard is available or acceptable, but one can be constructed, for example by an expert panel: calculate sensitivity and specificity against the constructed standard, describe it in the label, and create it independently of the new test’s results, ideally before specimens are collected.
  4. No reference standard is available and none can be constructed: report measures of agreement.

The guidance’s list of statistically inappropriate practices begins with this one: “Avoid use of the terms ‘sensitivity’ and ‘specificity’ to describe the comparison of a new test to a non-reference standard.”

FDA regulations show the distinction in action. The special controls for influenza antigen detection tests at 21 CFR 866.3328 set different criteria by comparator type. Against “an FDA-cleared nucleic acid based-test or other currently appropriate and FDA accepted comparator method other than correctly performed viral culture,” the requirement is stated as PPA (point estimate at least 80%, 95% lower bound at least 70%) and NPA (at least 95%, lower bound at least 90%). Against “correctly performed viral culture method,” the requirement is stated as sensitivity (influenza A at least 90% with lower bound at least 80%; influenza B at least 80% with lower bound at least 70%) and specificity (at least 95% with lower bound at least 90%).

The dual 510(k) and CLIA waiver guidance (2020) makes the same point for binary qualitative tests with an analytical cutoff: clinical sensitivity, specificity, likelihood ratios and predictive values can be estimated, but “when certain types of CMs are used in the study, measures such as positive percent agreement (PPA) and negative percent agreement (NPA) should be estimated instead,” referring to CLSI EP12 for details.

Can one study report both?

Yes, and it is a useful design when the field has both an established reference method and a modern cleared assay. One point-of-care molecular Strep A study on this site collected a second throat swab for both blood agar culture and an FDA-cleared molecular assay. Against culture it reported sensitivity of about 96% and specificity of about 97%. Against the molecular assay it reported PPA of about 94% and NPA of nearly 100%. The same device produced two pairs of numbers with different names because the comparators differ in status, and the lower PPA against the more sensitive molecular comparator is exactly what one would expect.

How should composite comparators be handled?

A composite comparator combines several methods by a pre-specified rule, such as “positive if at least two of three assays are positive.” Under FDA’s 2007 guidance a combination can serve as a reference standard if the combining algorithm is fixed as part of the standard. Two cases on this site illustrate how this is reported:

  • A point-of-care STI study used a composite comparator for chlamydia and gonorrhea and a patient infected status algorithm for trichomonas, each requiring at least two of three nucleic acid tests positive. Results were reported as PPA and NPA for CT and NG, and as sensitivity and specificity for TV.
  • A reader-based influenza A/B antigen study used a composite reference of two FDA-cleared molecular assays plus cell culture, with at least two of three positive defining a positive, and reported sensitivity and specificity.

Whether a given composite is accepted as a reference standard is a decision to settle with FDA before the study, not after. What FDA is explicit about is the limit: avoid comparing a new test “to the outcome of a testing algorithm that combines several comparative methods (non-reference standards), if the algorithm uses the outcome of the new test.” A follow-up test triggered by the new test’s result, or a tie broken by it, builds the new test into its own comparator, which the guidance says “will likely be biased in favor of the new test.”

A sequential algorithm is acceptable when the decision to run additional methods depends only on comparator results. The guidance calls that approach potentially “statistically reasonable.”

Why does FDA caution against discrepant analysis?

Because revising the two-by-two table on the basis of a resolver test inflates apparent performance in a way that cannot be quantified. Discrepant resolution retests only the specimens where the new test and comparator disagree and then reclassifies the comparator result to match the resolver. FDA’s 2007 guidance states: “You should not use outcomes that are altered or updated by discrepant resolution to estimate the sensitivity and specificity of a new test or agreement between a new test and a non-reference standard.”

The reasons it gives are specific:

  • When the original two results agree, the method assumes without evidence that both are correct. “Even when the new test and non-reference standard agree, they may both be wrong.”
  • Results can move only from disagreement cells to agreement cells, so apparent agreement can only rise. “In fact, using a coin flip as the resolver will also improve apparent agreement.”
  • The revised table’s columns no longer represent condition status or any single comparator.

FDA states that it “is not aware of any scientifically valid ways to estimate sensitivity and specificity by resolving only the discrepant results, even when the resolver is a reference standard.” To obtain unbiased estimates, the resolver must be a reference standard and “you must resolve at least a subset of the concordant subjects.” Resolving discrepancies by repeat testing with the new test or the comparator “does not provide any useful information about performance.”

None of this prohibits running a second method on discordant specimens. Several case studies on this site report, for example, how many apparent false negatives were also negative on a second high-sensitivity assay, presented as descriptive information next to an unrevised primary table. That is the defensible use.

A worked example: one table, calculated correctly and incorrectly

Consider a hypothetical study of 600 subjects comparing a new test with an FDA-cleared comparator assay that is not a reference standard.

Comparator +Comparator −Total
New test +112 (a)9 (b)121
New test −8 (c)471 (d)479
Total120480600

Agreement measures with two-sided 95% Wilson score intervals, computed in Python. (The 2007 guidance uses score intervals in its examples; we confirmed the Wilson formula reproduces the guidance’s published intervals exactly.)

MeasureCalculationEstimate95% CI
PPA (new / comparator)112/12093.3%87.4% to 96.6%
NPA (new / comparator)471/48098.1%96.5% to 99.0%
Overall percent agreement583/60097.2%95.5% to 98.2%

Direction matters. Computing the reverse, the proportion of new-test positives that are comparator positive, gives 112/121 = 92.6% (86.5% to 96.0%), and the proportion of new-test negatives that are comparator negative gives 471/479 = 98.3% (96.7% to 99.2%). These are different numbers from the same table. This is why FDA recommends “explicitly stating the calculation being performed” and labels its own examples “positive percent agreement (new / non ref. std.).”

Overall agreement hides the imbalance. FDA’s guidance illustrates this with its own tables: a test with 96.5% overall agreement had PPA of only 67.8% (40/59). Overall agreement is weighted toward whichever group is larger, usually negatives. FDA “discourages the stand-alone use of measures of overall agreement,” including Cohen’s kappa.

Now the discrepant-resolution version. Suppose the 17 discordant specimens are sent to a resolver assay. Of the 9 new-test-positive, comparator-negative specimens, 6 are resolver positive; of the 8 new-test-negative, comparator-positive specimens, 3 are resolver negative. Reclassifying those 9 comparator results produces a revised table:

MeasureOriginal table”Revised” tableChange
PPA112/120 = 93.3%118/123 = 95.9%+2.6 points
NPA471/480 = 98.1%474/477 = 99.4%+1.3 points
Overall agreement583/600 = 97.2%592/600 = 98.7%+1.5 points

Every measure improved, and none of the 583 concordant results was examined. FDA’s guidance says this kind of revised table should not be presented in the final analysis “because it may be very misleading.” The correct report is the original table, the measures in the first table above, and, if useful, a descriptive note on resolver findings for the discordant specimens.

How should results be reported?

FDA’s 2007 guidance lists what it recommends reporting. Gathered in one place:

  • The two-by-two table of new-test results against the reference standard or comparator.
  • A description of the comparator, how it was performed, and whether it is a reference standard.
  • The pair of measures (sensitivity and specificity, or PPA and NPA) with “two-sided 95 percent confidence intervals,” expressed “both as fractions (e.g., 490/500) and as percentages (e.g., 98.0%).”
  • The direction of agreement for PPA and NPA (new test relative to comparator).
  • A complete accounting of subjects: number planned, tested, used in the final analysis and omitted.
  • Ambiguous and equivocal results, stratified by comparator outcome, rather than silently excluded. One option the guidance describes is reporting performance twice, with equivocals counted as positive and then as negative.
  • Results by clinical site, testing site and relevant subgroups, with the intended-use population reported separately. The guidance notes that results from subjects outside the intended-use population “should not be labeled as ‘specificity.’”
  • No predictive values or likelihood ratios when the comparator is not a reference standard, since condition status is unknown.

For CLIA waiver applications, add the invalid-result reporting the CLIA waiver guidance recommends: initial and final invalid percentages per operator with 95% two-sided confidence intervals, excluded from the performance calculations, with a rationale for why the rate is clinically acceptable.

Common mistakes to avoid

  • Calling agreement “sensitivity” because the comparator is very good. A high-sensitivity cleared assay is still a non-reference standard unless FDA has agreed otherwise for your intended use.
  • Expecting high PPA to prove the new test is right. FDA notes that “two tests could agree well, but both have poor sensitivity and specificity,” and that disagreement does not mean the new test is wrong.
  • Letting the new test influence the comparator, through reflex testing, tie-breaking or discrepant reclassification.
  • Leading with overall agreement or kappa instead of the PPA and NPA pair.
  • Deciding the comparator after the data are in. The 2007 guidance recommends deciding before collecting the first specimen whether you will report accuracy or agreement, and consulting FDA early. The comparator also sets the denominator for PPA or sensitivity, and with it the number of positives the study must enroll.

How Studybox approaches this

We settle the comparator, its status and the reporting language in the protocol before enrollment, alongside the statistical analysis plan, and recommend confirming them with FDA through a Pre-Submission. Comparator specimens are collected and handled so that the comparator result is independent of the investigational result, and discordant testing, where used, is pre-specified as descriptive.

Our team has supported 55+ FDA regulatory clearances across 510(k), dual 510(k)/CLIA waiver, OTC and EUA pathways, and works exclusively on IVD studies. Prospective collection runs at our network of 100+ pre-qualified U.S. sites under a pre-approved IRB protocol. More on our 510(k) studies, CLIA waiver studies and dual submission studies, or contact us.

Studybox Research Guide October 5, 2026

Let's talk IVD research

Wherever you fit in the research process, Studybox is here to support you.

Sponsor or clinic: tell us about your study or your site and we'll follow up within one business day.