Studybox Research

Studybox Research Guide

Lay User Studies for OTC IVDs: Designing a Home-Use Diagnostic Study

An OTC lay user study has to show that untrained people, using only the box and its instructions, can collect an adequate sample, run the test, read the result correctly and understand what to do next, and that the test still performs to its claim when they do.

← All guides October 5, 2026 11 min read

Field notes

Guide

A lay user study for an over-the-counter (OTC) IVD shows that intended users with no medical or laboratory training can carry out the whole workflow with only the labeling that will ship with the product. That means collecting their own specimen, running the test, reading the result and understanding what it means. It also shows that clinical performance holds up when they do. FDA’s evidence usually has three parts: a usability or human factors evaluation of the tasks, a comprehension evaluation of the labeling and the results, and a clinical performance study run by lay users in a home or home-like setting against an appropriate comparator.

No single FDA guidance covers every OTC IVD. The requirements come from several documents, which this article draws on:

SourceWhat it contributes
Design Considerations for Devices Intended for Home Use (FDA, November 24, 2014)Who the home user is; environmental, physical, sensory and cognitive considerations; labeling and human factors expectations
Applying Human Factors and Usability Engineering to Medical Devices (FDA; originally 2016, current version issued August 3, 2026)Critical tasks, simulated-use validation, participant numbers and representativeness
21 CFR 866.3984, OTC test to detect SARS-CoV-2Binding special controls for one OTC device type, including lay-user specimen collection, a home-setting clinical study and label comprehension
EUA template for home-use COVID-19 tests (FDA, pandemic era)The most detailed public FDA text on lay-user usability study design
Self-Monitoring Blood Glucose Test Systems for Over-the-Counter Use (FDA, September 29, 2020)A worked example of an analyte-specific OTC user-performance study

A regulatory point is worth knowing up front. Under 42 CFR 493.15(b)(1), tests “cleared by FDA for home use” meet the CLIA criteria for waived tests. The evidence that makes an OTC test acceptable to FDA therefore also carries its waived status, which is one reason the lay-user evidence gets close scrutiny.

What does an OTC lay user study have to demonstrate?

It has to show four things: lay users can collect an adequate specimen, perform the test steps correctly, interpret every possible result correctly (including invalids), and understand the labeling well enough to act on the result. In parallel, the test must meet its clinical performance claim when lay users run it.

For OTC SARS-CoV-2 tests, these elements are codified as special controls in 21 CFR 866.3984:

  • The intended use may only include respiratory specimens “for which there are performance data that demonstrate lay users can collect specimens without health care provider supervision in home settings or similar environments.”
  • The clinical study must be prospective and multisite, from “geographically diverse locations”, with participants “representative of the intended use population”, and “performed in the intended use setting (e.g., at home or a home-like environment)”.
  • The study size must be large enough that “the lower bound of the two-sided 95 percent confidence interval of the positive percent agreement with the comparator must be greater than 70 percent”, with additional risk mitigations such as presumptive negative results and serial testing.
  • Risk-control documentation must cover “the entire testing procedure from sampling to result interpretation”, based on “usability studies, user label comprehension, and flex studies” as applicable.

That regulation applies only to OTC SARS-CoV-2 tests. Its structure, though, with lay collection, home-setting performance, comprehension and flex studies tied to risk, is a useful model for any OTC IVD. It also matches the general expectations in the home-use and human factors guidances. Flex studies are covered in a separate guide; this one deals with the human side.

Who counts as a lay user, and how do you recruit representative ones?

A lay user is someone without relevant medical or laboratory training who matches the intended-use population for the product. Recruiting means actively excluding people with healthcare or laboratory backgrounds, and deliberately spreading enrollment across age, education, literacy and the physical and sensory traits that could affect the tasks.

The 2014 home-use guidance defines a user as “a lay person such as a patient (care recipient), caregiver, or family member”, and “lay” as a person “without relevant specialized training”. It asks manufacturers to design for a range of “physical sizes, mobility, dexterity, coordination”, “vision and hearing abilities”, and “abilities to process information and literacy levels”. It also asks them to allow for the anxiety of someone dealing with a new diagnosis.

Recruitment criteria in FDA documents:

CriterionFDA textSource
Exclude trained people”Participants with prior medical or laboratory training or prior experience with self-collection or self-testing (including infectious disease home tests) should be excluded.”EUA home template
Spread of backgrounds”Enrollment population should represent different socioeconomic and educational backgrounds.”EUA home template
Age coverageEnrollment across bands from under 14 (parent or guardian testing a child) to 65 and over, with at least 30 children aged 2–13 across usability and clinical work for an all-ages claimEUA home template
Record demographicsAge range, education level, native language, “laboratory or healthcare work experience”, disease stateSMBG OTC guidance
Include naive users”At least 10% of the study participants should be naïve to SMBGs”SMBG OTC guidance
Functional limitationsPeople with a representative range of disease-related limitations (e.g., retinopathy, neuropathy) “included as representative users”HFE guidance §8.1.1
No employees; U.S. residentsEmployees “should not serve as test participants”; participants “should reside in the US”HFE guidance §8.1.1
Minimum per user group”In general, the minimum number of participants should be 15” per distinct user populationHFE guidance §8.1.1

Two of these interact. The HFE guidance treats groups that do different tasks as distinct user populations, so someone testing themselves and an adult testing a child count separately. That is why the EUA home template recommends “a minimum of 30 participants split evenly into two sections: 15 participants testing themselves and 15 participants testing another person.” If your labeling allows caregiver testing, plan for both groups from the start.

The EUA template adds a recruitment warning: “Recruitment by internet, especially if using monetary incentives, can drastically bias the population who enrolls in the study.” Recruiting through clinics where intended-use patients already present tends to give a more representative mix of education and testing experience. It also makes the screening question about healthcare employment easier to check.

Should lay users be observed or unobserved?

Both, at different stages. Usability evaluations are observed, because the point is to see and record every difficulty. The clinical performance study is run as close to real use as possible, with no help, no prompting, and no chance to watch anyone else.

DesignWhereObservationWhat it showsMain risk
Formative / usability studySimulated or actual use environmentObserved in person or by video; difficulties recordedWhich steps fail and why, before the design is lockedObserver intervention contaminating results
HF validation (simulated use)Realistic simulated environmentObserved, with no coachingThat critical tasks are performed without use errors that could cause harmTraining or prompting that real users won’t get
Clinical study, home-like siteStudy site set up to simulate a homeObserved for safety and data capture only; users isolated from one anotherPerformance against a comparator in lay handsUsers copying staff or other participants
Clinical study, at homeParticipant’s homeUnobserved, or remotely monitoredPerformance in the true use settingData completeness, comparator logistics

FDA’s documents line up with this sequence. The EUA template recommends that participants “be observed (either in person or by remote visual monitoring, such as a video conference) during sample collection and all difficulties should be noted.” It also says that “prior to conducting the at home clinical study, you should first complete an observed usability study.” For the clinical phase it asks that sites “be set up in a way that precludes a user from seeing or hearing other users performing the test.” The SMBG OTC guidance says that “no other training or prompting should be provided,” that subjects “should not receive assistance from a study technician,” and that they “should be sequestered”.

Can you combine usability with the clinical study? FDA allows it but warns that it “presents risk of a failed clinical study if there are problems with the instructions for use”. Combining saves a phase only if the instructions are already right. A separate, smaller usability round is the cheaper place to find out that they are not.

One subtle source of bias is the comparator sample. If staff collect a reference swab from the same site first, participants learn the technique by watching. The EUA template says to ensure participants “are not provided additional training by observing how healthcare providers collected a sample.” Randomizing swab order, or collecting the comparator after the lay-user test, deals with it.

How do you test result interpretation and label comprehension?

Show participants a set of results whose true answer you know, including positive, negative, invalid and faint or borderline lines, and ask them to read each one and say what they would do next. Separately, test whether they understood the key labeling concepts. Assess readability before the study starts.

The CLIA waiver guidance (written for untrained professional operators, but the method carries over) recommends that you “show various possible test results and control results that are positive, negative, and invalid, and ask the untrained operator to read these results.” For OTC tests, manufacturers commonly do this with contrived devices: cassettes or strips made to show a defined pattern, such as a faint test line, control-line-only, no control line, or a smeared window. Every participant then sees the hard cases, not only the clear results their own sample produced. The EUA template frames this as evaluating “interpreting positive, negative, and invalid results”, and notes “significant risks associated with misinterpretation and misuse of test results.” The SARS-CoV-2 special controls name “user label comprehension” as part of the risk-control evidence.

Readability is a pre-study check. FDA’s SMBG OTC guidance asks for a readability assessment and a reading level “at an 8th grade level or less”, using “the Flesch-Kincaid, SMOG, or equivalent”. The EUA home template and the CLIA waiver guidance both use “no higher than a 7th grade” level for quick-reference instructions. Write to the lower level if you can.

A worked example, and its limit. Suppose 30 participants each read 10 contrived devices, giving 300 reads, and 294 are correct (98.0%). The two-sided 95% Wilson interval is 95.7% to 99.1%. That interval treats the 300 reads as independent. They are not: a participant who confuses a faint line once will tend to do it again. A more honest analysis reports agreement for each pattern (especially faint positives and invalids) and per participant, and checks whether errors cluster in a few people. Ten reads of the faint-positive device from 30 people is really 30 observations of that pattern, not 300.

Zero errors in a small sample also means less than it appears. With 0 errors in 30 reads of the faint-positive pattern, the 95% Wilson upper bound on the error rate is still 11.4%. If faint positives are the highest-risk misread for your device, plan more reads of that pattern than of the easy ones.

How large does the clinical part of the study need to be?

That depends on the performance claim. Where FDA sets the claim as a lower confidence bound, as 21 CFR 866.3984 does (PPA lower bound greater than 70%), the number of comparator-positive subjects determines what is achievable. Using the Wilson score interval, the smallest number of agreeing positives needed to put the two-sided 95% lower bound above 70% is:

Comparator positives (n)Minimum agreeing (x)Observed PPAWilson 95% lower bound
302686.7%70.3%
403485.0%70.9%
504284.0%71.5%
604981.7%70.1%
806581.2%71.3%

The pattern matters for planning. With 30 positives, a test whose true PPA is around 80% will usually fail to clear the bound: 24/30 (80.0%) has a lower bound of 62.7%. With 80 positives, an observed PPA a little above 81% is enough. Calculate the positive count from the PPA you realistically expect, then work back to total enrollment using prevalence at your sites. The regulation fixes a two-sided 95% interval but not the method, so name the method (Wilson score, as here) in the statistical analysis plan, which the regulation requires to be predefined.

For a quantitative OTC test, the logic is the same but the endpoint is agreement within limits. The SMBG OTC guidance, for example, asks for “at least 350 different subjects” and for 95% of results within ±15% of the comparator and 99% within ±20%.

How does the lay user study relate to human factors validation?

They overlap but are not the same. Human factors validation shows, under simulated use, that users can perform the critical tasks without use errors that could cause serious harm. The clinical lay user study shows that, in real use, the results are accurate enough. A submission usually needs both, and failures in either lead back to the same place: the design and the instructions.

The HFE guidance defines a critical task as one which, “if performed incorrectly or not performed at all,” would or could cause serious harm. For an OTC IVD the critical tasks usually include specimen collection, the timed steps and reading the result. Two of its rules matter most for OTC diagnostics:

  • Fixing the label is not enough on its own. If validation shows use errors on critical tasks, “stating in a premarket submission that you mitigated the risks by modifying the instructions for use … is not acceptable unless you provide additional test data demonstrating that the modified elements were effective.”
  • Train participants only as real users would be trained. “If you anticipate that most or all users would receive minimal or no training, then the test participants … should not be trained.” For an OTC kit, that means the box, the insert, and any app or video that actually ships with it.

The 2014 home-use guidance adds that “labeling alone generally does not offer sufficient risk control for the home use environment because warning labels, especially lengthy ones, can be ignored by or confusing to the user.” Where the device can be designed so the error cannot happen (a fixed-volume dropper, a cassette that only accepts the swab one way, an app timer that hides the result until the read window), that design change carries more weight than a stronger warning.

What do real OTC and home-use studies look like?

Studybox case studies show how the same principles apply across specimen types and pathways:

  • An OTC 510(k) study for a COVID-19 antigen home test: lay users self-sampled and self-tested in a simulated home setting, with separate usability and readability studies covering interpretation of low-positive results.
  • An EUA study for an OTC home rapid antigen test and one for a non-prescription home molecular test, where lay users (and adults testing children) ran the full workflow from swab to result. The molecular study added an evaluation of whether lay users could correctly interpret near-cutoff samples.
  • Home-collection studies, where the user task is collecting and mailing a sample rather than testing. That shifts the emphasis to specimen adequacy and the shipping path.

How Studybox approaches this

Studybox Research works only on in vitro diagnostics, and its team has supported 55+ FDA regulatory clearances across 510(k), dual 510(k)/CLIA waiver, OTC and EUA pathways. For lay user studies, the practical constraints are recruitment that is truly lay and representative, and sites that can keep participants isolated while capturing comparator samples. Our 100+ pre-qualified U.S. sites are urgent care, physician office and point-of-care clinics where intended-use patients already present. Prospective specimen collection can run under a pre-approved IRB protocol and launch in as little as one week.

Typical study activation is about four weeks, against a 3–6 month industry average, which matters for seasonal analytes. If you are planning an OTC or home-use study, see our 510(k) studies overview or contact us with your intended use and labeling draft. We can review the usability, comprehension and clinical phases together before any of them is locked.

Studybox Research Guide October 5, 2026

Let's talk IVD research

Wherever you fit in the research process, Studybox is here to support you.

Sponsor or clinic: tell us about your study or your site and we'll follow up within one business day.