Evidence-Vetting Checklist
Evidence-Vetting Checklist
Use this checklist before publishing any learning module, claim, rubric descriptor, or controversial case study.
Gate 1: Claim Definition
- Claim is specific and testable.
- Claim scope (population, context, conditions) is explicitly stated.
- Claim type identified: descriptive, correlational, causal, normative.
Gate 2: Source Class and Quality
Source class hierarchy
Identify which class the supporting source belongs to and document it in the evidence log:
| Class | Description | Weight |
|---|---|---|
| Primary | Original peer-reviewed study, RCT, pre-registered trial | Highest |
| Secondary | Systematic review, meta-analysis, narrative review | High |
| Expert consensus | Clinical guideline, professional body position statement | Moderate |
| Tertiary | Textbook chapter, encyclopedia entry, established reference | Moderate |
| Grey literature | Government report, white paper, conference abstract | Low–Moderate |
| Commentary / opinion | Editorial, expert opinion without systematic review | Low |
Checklist:
- Source class identified and recorded in evidence log.
- Preference given to primary or secondary sources; lower-class sources noted as provisional.
- For meta-analyses: PRISMA checklist adherence noted; heterogeneity (I²) considered.
- For expert consensus: issuing body, date of consensus, and scope of agreement documented.
- Publication venue and author credentials checked.
- Funding and conflict-of-interest disclosure reviewed.
- At least one independent corroborating source identified for non-trivial claims.
Gate 3: Methodological Rigor
- Study design is appropriate to the claim type (e.g., RCT for efficacy, cohort for risk, qualitative for lived experience).
- Sample size is reported; assess whether it is sufficient for the claimed effect size (underpowered studies flagged).
- Control conditions are present where required; absence of a control group is noted as a limitation.
- Blinding and randomization procedures documented for experimental studies.
- Replication status assessed: single study, independent replications, or pre-registered replication.
- Sample and context are relevant to adult online learning use case; any extrapolation noted.
- Limitations section extracted and summarized in evidence log.
- Effect sizes or practical significance documented where available; statistical significance alone is insufficient.
Gate 4: Causality and Generalizability Limits
- Observational findings (cross-sectional, case study) are not written as causal claims.
- Confounders identified: demographic, environmental, or study-design factors that may explain the effect.
- Mediation and moderation analyses noted if present; absent analyses flagged where relevant.
- Animal or cell-culture findings are not generalized to human outcomes without explicit caveat.
- Correlations are presented with uncertainty language (e.g., "associated with", "linked to").
- Generalizability assessed: WEIRD bias (Western, Educated, Industrialized, Rich, Democratic samples), age, clinical vs. non-clinical populations.
- Single-study findings are not presented as established consensus.
Gate 5: Language Integrity
- Remove or qualify hype terms: breakthrough, cure, revolutionary, guaranteed, proven.
- Replace absolute statements ("always", "never", "everyone") with evidence-quality qualifiers.
- Distinguish recommendation from established evidence from emerging finding.
- Uncertainty language matches actual evidence strength (see evidence tier in Gate 8).
Gate 6: Citation Validity and Attribution
Citation validity
- In-text citations and reference list entries are consistent and complete.
- Citation points to the exact page, section, or data table supporting the claim.
- No citation laundering: secondary sources are not cited as if they are primary.
- Publication is current: for rapidly-evolving fields (e.g., neuroscience, AI, digital learning), sources older than 10 years are flagged; for stable fields (e.g., foundational cognitive psychology, established developmental theory), note recency status and confirm no superseding evidence exists.
- Retracted papers or superseded guidelines are not cited; checked against retraction databases if uncertain.
- AI-generated summaries are manually validated line-by-line against the original source before use.
Quoting, paraphrasing, and attribution
- Direct quotes use quotation marks and include page or paragraph number.
- Paraphrases are substantially reworded (not merely word-substituted) and retain the original meaning without distortion.
- Attribution is given for all specific claims, data points, figures, and tables.
- Images, diagrams, or adapted materials include source, creator, and license information.
- Permission or fair-use justification documented for copyrighted material.
Gate 7: Bias and Balance
- Counter-evidence was actively searched and documented.
- Competing interpretations represented fairly.
- Reviewer confirms no ideological loading language.
Gate 8: Publication Readiness
- Evidence quality tier assigned per claim:
Tier Astrong convergent evidence (multiple high-quality replications)Tier Bmoderate evidence with known limitations (mixed replication, methodological concerns)Tier Cprovisional guidance (single study, low-quality evidence, or expert opinion only)
- Reviewer sign-off completed.
- Escalation route followed for unresolved disputes (see Escalation and Rejection Policy below).
Red-Flag Reference Table
Use this table during content review. When a red-flag phrase or pattern is detected, replace it with the preferred alternative or add the required qualifier.
| Red Flag | Why It Is Problematic | Preferred Alternative |
|---|---|---|
| "Studies prove that…" | "Prove" implies certainty science cannot provide | "Research suggests…" / "Evidence indicates…" |
| "Revolutionary / groundbreaking / game-changing" | Hype language; overstates impact | Describe the specific finding without superlatives |
| "Guaranteed to…" | Implies certainty in outcomes | "Associated with improvement in…" |
| "Cures / eliminates / fixes" | Causal and absolute; rarely warranted | "May reduce…" / "Is linked to lower rates of…" |
| "All experts agree…" | Rarely true; suppresses legitimate disagreement | "A broad consensus suggests…" or cite specific bodies |
| "The science is settled" | Appropriate only for a narrow set of topics; misapplied as rhetorical closure | "Current evidence strongly supports…" |
| "As AI confirms…" or citing AI output as a source | AI output is not a primary source | Cite the underlying study or consensus document directly |
| "Multiple studies show…" without citations | Vague claim with no traceability | Cite specific studies; summarize convergence explicitly |
| "The only explanation is…" | Dismisses alternative interpretations | "One well-supported explanation is…" |
| "This approach transforms / supercharges…" | Marketing language; not scientific | Describe the measured outcome and effect size |
| "Proven brain hack / life hack" | Trivializes evidence and implies effortless results | Describe the mechanism and the conditions of effectiveness |
| "100% of participants…" without n and context | Absolute claim; small-sample artefact | Report the actual n and confidence interval |
| "Never / always" as factual statements | Overgeneralization from limited samples | Qualify: "In most cases…", "Evidence suggests…" |
| "Since time immemorial / ancient wisdom proves…" | Appeals to tradition; not evidence | Reference specific cultural practice and note evidence status |
| Uncited paragraph following an AI-generated summary | AI content laundered as authoritative | Add source citation; manually verify each factual claim |
Escalation and Rejection Policy
[HUMAN-REQUIRED] The thresholds and final policy below must be reviewed and approved by a qualified human reviewer before repository-wide adoption. The draft criteria below are provided for human review.
Escalation triggers
Any of the following requires escalation to the Method Reviewer and Approver before publication:
- A claim is safety-sensitive (e.g., mental health, medical, child welfare) and rests on Tier C evidence or below.
- A claim contradicts established clinical guidelines without a compelling methodological rationale.
- Evidence is contested: reviewers disagree on tier classification after one revision cycle.
- A conflict of interest in the source cannot be ruled out and the claim is non-trivial.
- A single pre-registration or replication failure contradicts a module claim.
- AI-generated content cannot be fully traced to primary sources after manual review.
- Counter-evidence with comparable methodological quality to the supporting evidence has been identified.
Rejection criteria
A claim must be rejected and may not be published in its current form if:
- No credible primary or secondary source can be located after a documented search.
- The only supporting sources are retracted, superseded, or conflict-of-interest-compromised.
- The claim implies therapy, diagnosis, or clinical intervention without clinical oversight.
- The language cannot be revised to accurately reflect the evidence without losing instructional utility.
- Safety reviewers have flagged the claim as potentially harmful without a viable mitigation.
Remediation path
| Outcome | Required action |
|---|---|
| Escalated, resolvable | Revise claim language to match evidence tier; re-review within 5 business days |
| Escalated, unresolvable | Archive claim with dissent note; do not publish; open follow-up research task |
| Rejected | Remove claim; document reason in evidence log; notify content author |
| Safety-flagged | Remove immediately; escalate to Safety Reviewer within 24 hours |
Governance integration
This policy integrates with:
- Peer Review SOP — reviewer roles, rounds, and severity scale
- Instructional Design Protocol — content objectivity and verification standards
- Shadowwork Safety Standard — safety escalation for emotionally intensive content
Required Artifacts
- Evidence log (claim → source → source class → quality tier → caveat)
- Reviewer notes and decision log
- Final claim language approved for publication
- Escalation or rejection record (if applicable)