Short Answer
An IVDR Annex XIII performance evaluation rests on three legs: scientific validity, analytical performance and clinical performance. In the clinical performance study, sensitivity and specificity are two separate endpoints and the sample size is calculated separately for each: with n = z² × p × (1 - p) / d², showing an expected 95% sensitivity to within ±3% takes about 203 diseased specimens; the same precision for specificity is calculated on its own, and in a prospective consecutive design the total sample is found by dividing by prevalence. The acceptance criterion is not the point estimate but the lower bound of the 95% confidence interval clearing a pre-written threshold. The reference standard must be independent of the test, discrepant analysis must not be used, and the report follows STARD 2015 and ISO 20916.
Serteser Danismanlik is run by a biomedical engineer (BME MSc) who developed a medical-AI medical device, published it in a peer-reviewed international journal, and led the methodology of four PROSPERO-registered systematic reviews. We deliver the design of the IVDR clinical performance study, its sample-size rationale, the statistical analysis plan and the statistical section of the performance evaluation report directly to manufacturers, and white-label, as a subcontractor, to RA consultancies and CROs. Ethics-committee and TITCK submission paperwork, site conduct and the principal investigator role are not our lane.
The most important way IVDR differs from MDR is that the evidence has a different name and a different structure: performance evaluation rather than clinical evaluation, a performance evaluation report (PER) rather than a CER, post-market performance follow-up (PMPF) rather than PMCF. Annex XIII Part A defines the three legs of that evaluation: scientific validity (the association of the analyte with the clinical condition), analytical performance (the test measuring the analyte correctly) and clinical performance (the test result relating to the clinical condition in the target population). This article is about the third leg, because that is where the statistical errors happen most and, with the transition timetable, where teams get stuck most.
MDCG 2022-2 sets out the general principles of clinical evidence for IVDs; ISO 20916:2019 defines good study practice for clinical performance studies using specimens from human subjects. Here we translate those two frameworks into design decisions and numbers.
What a clinical performance study is not
A clinical performance study is not a repeat of analytical validation. Analytical performance shows how accurately the test measures the analyte in known specimens (accuracy, precision, limit of detection, cross-reactivity). Clinical performance shows the test's ability to separate those with the condition from those without it in the real target population: sensitivity, specificity, positive and negative predictive values, likelihood ratios and, where needed, clinical utility.
The practical consequence of the confusion is this: 99% agreement obtained on contrived specimens in the laboratory is not clinical performance evidence. The target population is the population written into the intended purpose; the sample comes from there.
The reference standard and three design decisions
The quality of a clinical performance study is decided by three design decisions before any number:
- Reference standard: The method that establishes the presence of the condition independently of the test under evaluation. The reference cannot be the test itself or a composite that includes the test result; otherwise incorporation bias arises. Blinded reading of the reference is documented.
- Specimen selection: Consecutively or randomly selected specimens reflecting the population in the intended purpose. A case-control design built from known positives and healthy controls inflates sensitivity and specificity; if used, its justification and limitation are written in the report.
- Verification flow: Every specimen is assessed with the reference standard. A design where only test positives are verified against the reference produces verification bias and invalidates the specificity estimate.
These three decisions are written into the protocol before the sample-size calculation. Reclassifying discordant results afterwards with a third method and resolving only the discordants (discrepant analysis) is biased and not accepted; no reclassification that was not pre-defined in the protocol is performed.
Separate sample sizes for sensitivity and specificity
The most frequent error is calculating a single "n" and treating it as sufficient for both sensitivity and specificity. The two endpoints have different denominator populations: sensitivity is calculated from diseased specimens, specificity from non-diseased ones. The precision-based formula is applied separately to each:
n = z² × p × (1 - p) / d²
Example: The intended purpose claims 95% sensitivity and you want to show it to within ±3%.
n(diseased) = 1.96² × 0.95 × 0.05 / 0.03² = 3.8416 × 0.0475 / 0.0009 ≈ 203
If the expected specificity is 98% with a precision of ±2%:
n(non-diseased) = 1.96² × 0.98 × 0.02 / 0.02² = 3.8416 × 0.0196 / 0.0004 ≈ 189
In a prospective consecutive design with a 20% prevalence in the target population, reaching 203 diseased specimens means collecting about 1,015 specimens; that also yields 812 non-diseased specimens, more than enough for specificity. The decisive number is the one required for the lower-prevalence class.
As the expected value approaches 1 and the sample shrinks, this Wald approximation becomes optimistic; the analysis plan states that Wilson or Clopper-Pearson intervals will be used and the calculation is inflated somewhat accordingly.
Acceptance criterion: the lower bound, not the point estimate
In the performance evaluation report, "sensitivity was found to be 96%" is not a result on its own. The acceptance criterion is written into the protocol in this form: "The lower bound of the 95% confidence interval for sensitivity must exceed 90%." This form ties the sample-size calculation to the acceptance criterion: for the lower bound to clear the threshold, the gap between the expected value and the threshold must exceed the chosen precision.
The same logic is built for specificity and for positive and negative predictive values. Because predictive values depend on prevalence, the report gives them both at the study prevalence and at the prevalence expected under the intended purpose.
To have the protocol and sample-size rationale of your IVDR clinical performance study reviewed from a methods perspective before Notified Body submission, request a free 15-minute scoping call.
Reporting: STARD 2015 and ISO 20916
The clinical performance study report follows the structure of ISO 20916 and is completed with the STARD 2015 checklist for diagnostic accuracy reporting: participant flow diagram, reference standard definition and blinding, handling of indeterminate and invalid results, the 2x2 table, all estimates with confidence intervals, and pre-defined subgroup analyses. Invalid and indeterminate results are not brushed aside as "excluded"; their rates are reported and their effect on the intended purpose discussed.
In the PER, the clinical performance section is linked to the same intended-purpose sentence as the scientific validity and analytical performance sections. The three legs support one another; strong clinical performance does not compensate for scientific validity waved through on weak literature. The principles in Notified Body findings on the CER literature search apply to the literature leg of an IVD as well.
Common Mistakes
- A prospective claim on a case-control sample: Sensitivity and specificity obtained from known positives and healthy controls cannot be carried over to a screening or primary-care intended purpose; spectrum bias is discussed in the report.
- One n for two endpoints: Separate samples for sensitivity and specificity, with the lower-prevalence class decisive.
- Including the test in the reference standard: If the result of the test under evaluation sits inside a composite reference, incorporation bias arises.
- Choosing the cut-off on the study data: The decision threshold is locked during analytical development; it is not optimized on the clinical performance study data.
- Waiting for the transition deadline: With IVDR class transition dates approaching, leaving the study to the last month is the most expensive mistake because of ethics-committee timelines and specimen collection time.
Related Articles
- Sample Size in a PMCF Study
- Notified Body Findings on the CER Literature Search
- How to Design a SaMD Clinical Validation Study
- How to Write the CER Statistics Section
In an IVDR clinical performance study the number is the consequence of the design: if the reference standard is independent, the sample comes from the target population and the acceptance criterion is written on the lower confidence bound, the sample-size calculation is a one-line job. Otherwise no sample size rescues the study. The experience behind the method is on the about page.
Scope, timeline and budget differ for every file; we settle them in a free scoping call: