Short Answer
Sample size in a PMCF study is calculated for precision, not for a hypothesis test: you first write down which safety or performance rate you want to show, and with what confidence-interval width. For a proportion the formula is n = z² p(1-p) / d²; estimating an expected 5% complication rate to within ±2% takes roughly 460 patients. If no events are expected, the rule of three applies: zero events in n patients puts the 95% upper bound at about 3/n, so showing a rate below 1% needs at least 300 patients. For survival-type endpoints the number of events, not the number of patients, is decisive. The MDCG 2020-7 plan template asks for this rationale; acceptance criteria are tied to the claims in the CER and the residual risks in ISO 14971.
Serteser Danismanlik is run by a biomedical engineer (BME MSc) who developed a medical-AI medical device, published it in a peer-reviewed international journal, and led the methodology of four PROSPERO-registered systematic reviews. We deliver the sample-size rationale of the PMCF plan, the statistical analysis plan and the statistical section of the MDCG 2020-8 evaluation report directly to manufacturers, and white-label, as a subcontractor, to RA consultancies and CROs. Site conduct, monitoring and the principal investigator role are not our lane; we build the method behind the number.
The weakest sentence in most PMCF plans is "approximately 100 patients will be followed." When a Notified Body assessor reads it, the question is not "why 100" but "what will you be able to show with 100 patients." The MDCG 2020-7 PMCF plan template explicitly expects a rationale for the chosen method and the sample size; the MDCG 2020-8 evaluation report template expects the results to be compared against pre-defined acceptance criteria. Without a rationale, any number is questioned.
This article covers the logic of PMCF sample-size calculation, formulas and worked numbers for four typical scenarios, and how acceptance criteria are linked to the CER and the risk management file. The aim is not a statistics lesson but a rationale you can write into the plan.
First the question: what is PMCF trying to confirm
MDR Annex XIV Part B defines the purpose of PMCF as confirming the safety and performance of the device throughout its expected lifetime, identifying previously unknown side effects and contraindications, monitoring the acceptability of identified risks, and detecting off-label use. The sample-size calculation depends on which of these purposes is primary for the plan.
In practice there are three types of question:
- Estimating a rate: "What is the one-year revision rate" or "what is the rate of device-related serious adverse events." A precision-based calculation.
- Ruling out a rare event: The claim "this event occurs in fewer than 1%." Situations with zero or very few expected events; the rule of three and the Poisson approximation.
- Time-dependent endpoints: Implant survival, time to reintervention. The number of events and the follow-up time, not the number of patients.
Whichever question the evidence gaps flagged in the clinical evaluation correspond to, the sample is built for that question. The CER statistics section article is the foundation for defining those gaps.
Scenario 1: Estimating a rate to a given precision
You will estimate a safety or performance rate with a 95% confidence interval. The formula:
n = z² × p × (1 - p) / d²
Here z is 1.96 for 95%, p is the expected rate, and d is the desired half-width (the ± part of the interval).
Example: The expected one-year complication rate from the literature and the CER is 5%. You want to show it to within ±2% (between 3% and 7%).
n = 1.96² × 0.05 × 0.95 / 0.02² = 3.8416 × 0.0475 / 0.0004 ≈ 456
So roughly 460 evaluable patients. Add 10 to 15% for loss to follow-up. Showing the same rate to within ±3% takes 203 patients; to within ±1%, 1,825. Doubling the precision quadruples the sample; the plan should state this trade-off explicitly.
For small rates and small samples this Wald approximation is optimistic; the SAP states that Wilson or Clopper-Pearson intervals will be used in the analysis.
Scenario 2: Zero events and the rule of three
Some PMCF questions take the form "this event is not expected; show that it does not occur." If no event is observed in n patients, the 95% upper confidence bound on the true rate is approximately 3/n. This is the rule of three.
- Zero events in 100 patients: upper bound about 3%.
- Zero events in 300 patients: upper bound about 1%.
- Zero events in 1,000 patients: upper bound about 0.3%.
So supporting the claim "this event occurs in fewer than 1%" with zero events needs at least 300 patients. If one or two events are observed the upper bound rises; the plan should anticipate that case too. The threshold is written into the plan like this: "If no serious device-related event is observed in n = 300 patients, the 95% upper confidence bound stays below 1% and residual risk R-07 is considered acceptable."
Scenario 3: Survival-type endpoints
For endpoints such as implant survival, the device remaining functional, or time to reintervention, what matters is not the number of patients but the number of observed events and the follow-up time. The precision of a Kaplan-Meier estimate at a given time point depends on how many patients remain at risk at that point and how many events have been observed.
The plan states: the target time point (for example 2 years), the expected event rate, the acceptable confidence-interval width, the expected loss to follow-up, and the number of patients to enrol at the start as a result. The same precision cannot be obtained by shortening follow-up and enrolling more patients; the number of patients reaching the time point is decisive.
To have your PMCF sample-size rationale and analysis plan reviewed from the Notified Body's perspective, request a free 15-minute scoping call.
Scenario 4: Registries and real-world data
In PMCF run from a registry or routine clinical data, the sample question shifts from "how many patients to enrol" to "which patients are included and is the precision sufficient in subgroups." A large registry gives a narrow overall interval, but if the critical subgroup (a specific indication or size, for instance) is small, a separate calculation is needed for that subgroup. The plan writes the primary analysis population, the subgroups, and the expected precision for each. See real-world evidence protocol and SAP for the design principles.
Where the acceptance criteria are anchored
The second half of the sample-size rationale is the acceptance criterion. "Results will be evaluated" is not a criterion. Criteria come from three sources:
- Claims in the CER: Performance and safety thresholds defined against the state of the art in the clinical evaluation. The PMCF result is compared with these thresholds.
- ISO 14971 residual risks: Every residual risk judged acceptable in the risk management file has an upper bound on its occurrence rate; PMCF shows that bound is not exceeded.
- Previous PMCF and PSUR findings: Comparison with the previous period and trend analysis in the update cycle.
Each acceptance criterion is written as a measurable sentence in the plan and answered with the same sentence in the MDCG 2020-8 report. What happens when a criterion is not met (CER update, risk-file revision, assessment of field corrective action) also goes into the plan.
Common Mistakes
- Counting a user survey as clinical data: Survey-based PMCF is weak for estimating safety rates because of response bias and unverified event reporting; it can be supporting, not primary.
- Confusing precision with power: In PMCF there is usually no hypothesis under test; instead of a power calculation, write which interval width is acceptable.
- Ignoring loss to follow-up: The evaluable number is the target; the enrolment number is inflated by the expected loss.
- Settling for a single overall rate: If residual risks differ by subgroup (indication, size, user experience), subgroup precision is justified separately in the plan.
Related Articles
- Notified Body Findings on the CER Literature Search
- How to Write the CER Statistics Section
- IVDR Clinical Performance Study: Sensitivity, Specificity and Sample Size
- Real-World Evidence: Protocol and SAP
A good sample-size rationale in a PMCF plan is two sentences: what we will show and to what precision, and which acceptance criterion that number is tied to. Once those two sentences are written, the number itself stops being the subject of debate. The experience behind the method is on the about page.
Scope, timeline and budget differ for every file; we settle them in a free scoping call: