Draft revision of the Institute’s certainty assessment… — submissions
The 15 submissions received, published in full with declared interests and secretariat responses.
§2Submissions and responses
15 submissions were received. Each is published in full below with its declared interest, the secretariat response and the disposition. The Institute publishes submissions it did not accept in the same form as those it did.
Quantitative claims are reproduced without the method that produced them
This submission addresses version 2 from the standpoint of a reader who will encounter its output rather than its text.
Several figures in the draft are quoted from sources that determined them by different methods. A figure obtained by one determination and a figure obtained by another are not comparable, and the draft places them in the same sentence without distinguishing them.
The respondent, an analytical chemist, proposes that every quantitative claim carry the method that produced it at the point of use rather than in the reference.
The secretariat accepts this submission. Placing two figures side by side is an implicit claim that they are the same kind of quantity, and in the cases identified they were not.
Every quantitative claim now carries the determination that produced it at the point of use, and figures obtained by non-comparable methods are no longer presented in the same row or sentence.
The framework should say when a rating must not be issued at all
The respondent notes that version 2 has to work where the evidence base is very thin as well as where it is deep, and submits with that in view.
A certainty rating over an empty evidence base is a rating of nothing, and the machinery of domains and downgrades applied to no studies produces a number that looks like an assessment. The framework should forbid this rather than leave it to judgement.
The respondent proposes an explicit rule: where no eligible study exists for an outcome, no rating is issued and the outcome is recorded as unassessed.
The secretariat accepts this submission without qualification. It identifies a defect the Institute had not corrected.
Where no eligible study exists, no certainty rating is issued, no domain assessment is rendered, and the outcome is recorded as not assessed with the reason stated. The rule is applied retrospectively across the series.
The database set omits sources relevant to the compounds in scope
The respondent’s comment on version 2 arises from having applied a comparable framework to the same compounds.
The respondent, an information specialist, states that regional databases index trials of several compounds in the Institute's scope that are not indexed in the listed sources.
The respondent proposes that the named regional databases be added to the standard set.
The secretariat accepts this submission in part. Two of the named sources are added to the standard set. The remainder are added as conditional sources searched where the review question concerns a compound developed or registered in the relevant region.
The standard database set is extended, the conditional sources are named together with the trigger for searching them, and every review states which sources were searched and which were not, with the reason.
Point estimates are given without an interval
This submission concerns version 2 and makes one point.
Several estimates in the draft appear as single figures. The respondent states that a point estimate without an interval invites a precision the underlying data do not support, and that the effect is worst where the estimate is drawn from a small contributing set.
The respondent proposes that no point estimate appear anywhere in the document set without its interval, including in summary tables and in the abstract.
The secretariat accepts this submission in part. Intervals are added wherever the source reports one. The proposal is declined for figures the source published without an interval, because the Institute will not compute an interval a source did not report.
Every estimate now carries its interval where the source reported one, and where it did not, the estimate is annotated as reported without an interval rather than left to appear as a precise figure.
Sponsor concentration should be a recorded characteristic of a body of evidence
Having read the draft of version 2, the respondent puts one point to the committee.
Where every contributing trial has one sponsor, the usual protections of replication do not apply: the same protocol, the same analytical conventions and the same investigators recur, and an error common to them is invisible.
The respondent proposes that sponsor concentration be a recorded characteristic, without prescribing a downgrade.
The secretariat accepts recording and declines an automatic downgrade.
Sponsor concentration is now recorded for every body of evidence. It is not an automatic downgrade, since in this field a concentrated sponsor set is often the only set that exists, and a rule producing a uniform downgrade would carry no information.
The document should describe how an assessment becomes a decision
This is a submission on version 2.
The respondent states that the methodology stops at the certainty rating and that the step from a rating to a decision is where most disagreement actually occurs.
The respondent proposes an evidence-to-decision framework.
The respondent notes that submission 005 has already been made and confines this submission to a matter not covered by it.
The secretariat notes this submission and records the point as correct about the field and not about this document.
No amendment arises. The Institute assesses evidence and does not make decisions, because a decision requires values and a resource context the Institute does not hold. The boundary is stated on the methodology page and the submission has prompted it to be stated at the head of the document rather than in the section on scope.
It is not clear which provisions bind the assessment committee
This submission concerns the draft of version 2. The respondent assesses evidence for a national body and the observation arises from applying comparable guidance.
The respondent states that the document mixes requirements with descriptions of current practice in the same voice, so that a reader cannot tell which departures would be a breach and which would be a change of habit.
The respondent proposes that binding provisions be distinguished typographically and listed.
The secretariat accepts this submission. A rule indistinguishable from a description is not enforceable and does not reassure.
Binding provisions are now stated in a fixed form, are listed together in an annex, and a departure from any of them must be recorded in the document it affects with the reason, while descriptive passages are marked as descriptions of practice.
The framework does not say when a concern warrants one level and when it warrants two
The respondent has read version 2 in draft and makes one submission.
Serious and very serious are both available in every domain and the framework gives no criterion for choosing between them. In practice the choice determines the published rating more often than the choice of domain does.
The respondent proposes worked criteria for the two-level downgrade in each domain.
The secretariat accepts the proposal for three domains and declines it for the remainder.
Worked criteria are now published for imprecision, inconsistency and indirectness. For risk of bias and publication bias the judgement is left to the assessors with a requirement to state the reason, because a criterion expressed in advance would be applied to cases it was not written for.
Nothing states how many assessors rate a body of evidence
The respondent read version 2 in draft. The point applies to it and to the series generally.
A rating produced by one person and a rating produced by two independently with adjudication are different objects, and the framework does not distinguish them.
The respondent proposes that the number of assessors and the adjudication method be recorded on every rating.
The secretariat accepts this submission.
Ratings are made independently by two assessors and adjudicated by a third where they differ. The method and the fact of any adjudication are now recorded with the rating.
Nothing prevents a conclusion stronger than its certainty rating
This is a submission on version 2, made from a statistical standpoint.
The respondent states that the framework rates certainty and then leaves conclusion wording to the author, so that a low rating and a confident conclusion can coexist in one document.
The respondent proposes that conclusion wording be drawn from a fixed set tied to the rating.
The secretariat accepts this submission. The coupling is the mechanism by which a rating changes what a reader takes away.
Conclusion wording is now drawn from a fixed set of formulations tied to the certainty rating, the mapping is published in the methodology document, and a document whose conclusion does not match its rating cannot be ratified.
The framework does not say what new evidence would change a rating
The draft of version 2 was read by a respondent whose concern is its interoperability with published certainty guidance.
A rating is a judgement about a body of evidence at a date. Without a statement of what would change it, a reader cannot tell whether a newly published trial is material.
The respondent proposes that every rating carry a statement of the finding that would alter it.
The secretariat accepts this submission.
Every rating now carries a statement of what would change it, and the search date against which it was made. A trial meeting the stated description triggers reassessment ahead of the ordinary cycle.
A composite outcome is rated as one outcome
The respondent submits on version 2, on a matter that is not specific to this draft but is visible in it.
A composite combines components of unequal importance and often of unequal effect. Rating it as a single outcome hides both, and the framework does not require the components to be reported.
The respondent proposes that where a composite is assessed, the components be reported and their certainty rated separately.
The secretariat accepts this submission.
A composite outcome is now reported with its components, each rated separately, and the rating of the composite states which component drove it.
The document is unreadable without specialist training
The respondent submits on version 2. A framework of this kind is judged by whether two competent assessors applying it to the same evidence reach the same rating.
The respondent states that the draft is written for a reader who already understands certainty grading, and that the people most affected by the subject matter will not reach the assessment at all.
The respondent proposes a plain-language summary at the head of every document, written to the same standard of accuracy as the document itself and not as a promotional abstract.
The respondent read submission 006 after drafting this one and has not altered it, the two points being distinct.
The secretariat accepts this submission in part. A plain-language summary is added. The proposal that it replace the technical abstract is declined, because the abstract is the part of the document other assessors read and cite.
Every document now opens with a plain-language summary of not more than 150 words, placed above the technical abstract and carrying the same certainty language, so that the two cannot diverge.
The treatment of imprecision does not work where the event did not occur
The respondent has read version 2 in draft and makes a single submission.
The imprecision rules in the draft operate on the width of an interval. For a harm observed zero times, there is no interval of the kind the rules assume, and the draft gives no instruction.
The respondent proposes that the rule for a zero-event outcome be stated explicitly and that it produce a rating reflecting what the exposure could have detected.
The secretariat accepts this submission. The gap would have been filled by ad hoc judgement, which is what a framework exists to prevent.
The framework now states the rule for zero-event outcomes: the rating is determined by the total exposure and the frequency that exposure could have detected, and the resulting statement records that the evidence is uninformative rather than reassuring.
No information-size criterion is applied before an interval is judged
The respondent read version 2 in draft and has confined this submission to a single provision.
A narrow interval obtained from a small number of events may still be the product of chance, and the framework as drafted would rate it precise. An information-size criterion is the standard remedy.
The respondent proposes that a criterion be applied before precision is judged.
The secretariat accepts the principle and applies it as a check rather than as a rule.
An information-size check is now applied and its result recorded. It is not applied mechanically, because a criterion computed on an assumed effect size can itself be the more arbitrary judgement, and the domain note states which consideration governed.
References cited on this page
References are numbered in order of first citation in this document. Each superscript in the text links to its entry below.
- International Organization for Standardization. ISO/IEC 17025:2017 General Requirements for the Competence of Testing and Calibration Laboratories. ISO/IEC Standard 2017;3rd edition. identifier not held by the Institute
Identifiers are reproduced only where the Institute holds them. Where a digital object identifier or PubMed identifier is not shown, the Institute has recorded the journal and year and has not constructed an identifier.