Clinical Evidence and Regulation

Literature Appraisal in Clinical Evaluation: MEDDEV 2.7/1 Appendix A6

September 16, 2026 · 8 min read · Burak Serteser

Short Answer

Literature appraisal is the scoring of every study included in a CER against pre-defined criteria, and the documentation of how much each one contributes to the conclusion. MEDDEV 2.7/1 Rev 4 Appendix A6 defines two sets of criteria: suitability (D1 appropriate device, D2 appropriate device application, D3 appropriate patient group, D4 acceptable report and data collation) and data contribution (D5 data source type, D6 outcome measures, D7 follow-up, D8 statistical significance, D9 clinical significance). The guidance does not impose a fixed scoring scale; the manufacturer pre-defines the scheme in the clinical evaluation plan and applies it consistently. The evidence hierarchy in MDCG 2020-6 Appendix III and bias tools such as RoB 2, ROBINS-I and QUADAS-2 are embedded in that scheme. The Notified Body cares less about the scores themselves than about the scheme existing in advance and being applied to every study the same way.

Serteser Danismanlik is run by a biomedical engineer (BME MSc) who developed a medical-AI medical device, published it in a peer-reviewed international journal, and led the methodology of four PROSPERO-registered systematic reviews. We deliver the design of the appraisal scheme, the per-study scoring table and the synthesis section as part of the CER literature and statistical core, directly to manufacturers and white-label, as a subcontractor, to RA consultancies and CROs; clinical evaluation sign-off stays with your named clinical evaluator.

The search may have been done correctly, the flow diagram built, the included studies listed; the assessor's next question still comes: which of these studies did you trust, how much, and why? Appraisal is the answer. MEDDEV 2.7/1 Rev 4 Appendix A6 defines the criteria for appraisal but does not impose a scoring scale, and that flexibility is where CERs most often become internally inconsistent.

This article covers the two sets of criteria in Appendix A6, how they are combined with the MDCG 2020-6 evidence hierarchy, where the bias tools from the systematic-review world fit, and how a weighting scheme that a Notified Body accepts is built.

Why appraisal is a separate stage

In the stages of MEDDEV 2.7/1 Rev 4, appraisal (Stage 2) sits between data identification (Stage 1) and analysis (Stage 3). Its purpose is to determine whether each identified piece of data is suitable for the clinical evaluation of this device and, if so, how much weight it carries. The distinction matters: when analysis is started without appraisal, a case report and a multicentre study are cited in the same sentence with the same weight, and the assessor questions the whole synthesis.

Appendix A6.1 also lists examples of studies lacking scientific validity: publications where the device is not identified, inadequate patient numbers, improper statistical methods, abstract-only information, sponsored reports without disclosure of conflicts of interest. These examples are a draft of the exclusion criteria.

Suitability criteria: D1 to D4

Suitability asks whether the study is relevant to this device's clinical evaluation at all:

  • D1, appropriate device: Did the study examine the device under evaluation, a demonstrated equivalent, or merely a similar device? Directly linked to the equivalence table; data on a device whose equivalence has not been shown falls here.
  • D2, appropriate device application: Was the device used as described in the intended purpose, or off-label or with a different technique?
  • D3, appropriate patient group: Does the study population reflect the population in the intended purpose (age, disease stage, comorbidity)?
  • D4, acceptable report and data collation: Are methods and results reported in enough detail to allow assessment?

A simple, pre-defined scale is used for each criterion (for example fully suitable, partly suitable, not suitable); when "partly" is chosen, the reason is written. An unsuitable study is not discarded, but it is used as context rather than as evidence for the device.

Data-contribution criteria: D5 to D9

Data contribution asks how much strength a suitable study adds to the conclusion:

  • D5, data source type: Randomized controlled trial, prospective cohort, retrospective series, case report, registry data, vigilance record. The row where the design hierarchy is reflected.
  • D6, outcome measures: Does the endpoint the study measured match the device's clinical claim? Surrogate or clinical endpoint?
  • D7, follow-up: Is follow-up long enough to observe the endpoint and detect residual risks?
  • D8, statistical significance: Are effect estimates given with confidence intervals, is the sample adequate for the estimate, is the analysis method appropriate?
  • D9, clinical significance: Is a statistically significant finding also clinically meaningful; does the effect size make a difference to the patient?

Separating D8 and D9 is deliberate: a large study can make a small, clinically trivial difference significant; a small study can report a large but uncertain effect. The appraisal table keeps the two in separate columns.

The MDCG 2020-6 hierarchy and bias tools

MDCG 2020-6, in defining sufficient clinical evidence for legacy devices, presents an evidence hierarchy in its Appendix III: high-quality device-specific clinical investigations at the top, weaker designs below, and sources such as complaints and vigilance data at the bottom. This hierarchy can be used directly to build the scale for the D5 row and is a familiar reference for the Notified Body.

The bias tools of the systematic-review world provide the depth of D4 and D8: Cochrane RoB 2 for randomized trials, ROBINS-I for non-randomized studies of interventions, QUADAS-2 for diagnostic accuracy studies, PROBAST and PROBAST-AI for prediction and artificial-intelligence models. These tools do not produce the score; they document the reason for it. When you give a study a low contribution, "high risk in the RoB 2 randomization domain" is far more defensible than "weak study".

To have your CER appraisal scheme and per-study scoring table reviewed from the Notified Body's perspective, request a free 15-minute scoping call.

How the weighting scheme is built

Because Appendix A6 does not impose a scoring scale, the scheme is the manufacturer's responsibility. Accepted schemes share these features:

  • Pre-defined: The scheme is written in the clinical evaluation plan before the search is run. The scale is not changed after the data has been seen.
  • Simple and reproducible: A three- or four-point scale for the nine criteria; a total score or a categorical result (high, medium, low contribution). Complex weighted formulas do not read as transparent to an assessor.
  • Two axes reported separately: Suitability and contribution are not collapsed into one score; a suitable but weak study and a strong but unsuitable study are handled differently.
  • Effect on the conclusion discussed: Does removing low-contribution studies from the synthesis change the result? That sensitivity discussion belongs in Stage 3.
  • Dual assessment: Scoring is done by two people and a disagreement record is kept; in small teams the second assessor can be an external named methodologist.

The scheme itself goes into the CER annex, as does the completed table for every included study. The body text discusses only the summary and the effect on the conclusion.

Appraisal for artificial-intelligence devices

For software and AI-based devices, D6 to D8 need particular attention: the reported metric (AUC, sensitivity, Dice) may not match the clinical endpoint in the intended purpose; whether the test data was independent of the training data may not be reported; without external validation the contribution is low. The TRIPOD-AI and PROBAST-AI checklists are used to build the scale for these rows. See TRIPOD-AI and PROBAST-AI reporting for the detail.

Common Mistakes

  • Building the scale after seeing the studies: The most frequent finding. If the scheme is not in the plan, all scoring is treated as post hoc.
  • Using publication year or journal impact factor as a criterion: No such criterion exists in Appendix A6; the prestige of the journal does not show the suitability or contribution of the study.
  • Including an unsuitable study "with a low score": A study found unsuitable on D1 to D3 does not enter the table as device evidence; it moves to the context section.
  • Single-person scoring: The assessor asks about consistency; two assessors and a disagreement record close that question.
  • Forgetting the scores in the synthesis: If the appraisal table is built and Stage 3 still cites every study equally, the table is pointless.

Related Articles

Appraisal is the section that tells the reader how solid the CER's evidence base is, and the assessor reads that from the consistency of the scheme, not from the size of the scores. Nine criteria, a scale written in advance, two assessors and a completed table: when those four are in place, the section does not come back. The experience behind the method is on the about page and in the systematic review service.

Scope, timeline and budget differ for every file; we settle them in a free scoping call:

Free Scoping Call

Next step

Let's talk about your project.

In a free 15-minute intro call we listen to what you need and tell you which service tier fits.

Related service: CER Literature and Statistical Core