Skip to main content

Evidence-Vetting Checklist

Evidence-Vetting Checklist

Use this checklist before publishing any learning module, claim, rubric descriptor, or controversial case study.

Gate 1: Claim Definition

  • Claim is specific and testable.
  • Claim scope (population, context, conditions) is explicitly stated.
  • Claim type identified: descriptive, correlational, causal, normative.

Gate 2: Source Class and Quality

Source class hierarchy

Identify which class the supporting source belongs to and document it in the evidence log:

ClassDescriptionWeight
PrimaryOriginal peer-reviewed study, RCT, pre-registered trialHighest
SecondarySystematic review, meta-analysis, narrative reviewHigh
Expert consensusClinical guideline, professional body position statementModerate
TertiaryTextbook chapter, encyclopedia entry, established referenceModerate
Grey literatureGovernment report, white paper, conference abstractLow–Moderate
Commentary / opinionEditorial, expert opinion without systematic reviewLow

Checklist:

  • Source class identified and recorded in evidence log.
  • Preference given to primary or secondary sources; lower-class sources noted as provisional.
  • For meta-analyses: PRISMA checklist adherence noted; heterogeneity (I²) considered.
  • For expert consensus: issuing body, date of consensus, and scope of agreement documented.
  • Publication venue and author credentials checked.
  • Funding and conflict-of-interest disclosure reviewed.
  • At least one independent corroborating source identified for non-trivial claims.

Gate 3: Methodological Rigor

  • Study design is appropriate to the claim type (e.g., RCT for efficacy, cohort for risk, qualitative for lived experience).
  • Sample size is reported; assess whether it is sufficient for the claimed effect size (underpowered studies flagged).
  • Control conditions are present where required; absence of a control group is noted as a limitation.
  • Blinding and randomization procedures documented for experimental studies.
  • Replication status assessed: single study, independent replications, or pre-registered replication.
  • Sample and context are relevant to adult online learning use case; any extrapolation noted.
  • Limitations section extracted and summarized in evidence log.
  • Effect sizes or practical significance documented where available; statistical significance alone is insufficient.

Gate 4: Causality and Generalizability Limits

  • Observational findings (cross-sectional, case study) are not written as causal claims.
  • Confounders identified: demographic, environmental, or study-design factors that may explain the effect.
  • Mediation and moderation analyses noted if present; absent analyses flagged where relevant.
  • Animal or cell-culture findings are not generalized to human outcomes without explicit caveat.
  • Correlations are presented with uncertainty language (e.g., "associated with", "linked to").
  • Generalizability assessed: WEIRD bias (Western, Educated, Industrialized, Rich, Democratic samples), age, clinical vs. non-clinical populations.
  • Single-study findings are not presented as established consensus.

Gate 5: Language Integrity

  • Remove or qualify hype terms: breakthrough, cure, revolutionary, guaranteed, proven.
  • Replace absolute statements ("always", "never", "everyone") with evidence-quality qualifiers.
  • Distinguish recommendation from established evidence from emerging finding.
  • Uncertainty language matches actual evidence strength (see evidence tier in Gate 8).

Gate 6: Citation Validity and Attribution

Citation validity

  • In-text citations and reference list entries are consistent and complete.
  • Citation points to the exact page, section, or data table supporting the claim.
  • No citation laundering: secondary sources are not cited as if they are primary.
  • Publication is current: for rapidly-evolving fields (e.g., neuroscience, AI, digital learning), sources older than 10 years are flagged; for stable fields (e.g., foundational cognitive psychology, established developmental theory), note recency status and confirm no superseding evidence exists.
  • Retracted papers or superseded guidelines are not cited; checked against retraction databases if uncertain.
  • AI-generated summaries are manually validated line-by-line against the original source before use.

Quoting, paraphrasing, and attribution

  • Direct quotes use quotation marks and include page or paragraph number.
  • Paraphrases are substantially reworded (not merely word-substituted) and retain the original meaning without distortion.
  • Attribution is given for all specific claims, data points, figures, and tables.
  • Images, diagrams, or adapted materials include source, creator, and license information.
  • Permission or fair-use justification documented for copyrighted material.

Gate 7: Bias and Balance

  • Counter-evidence was actively searched and documented.
  • Competing interpretations represented fairly.
  • Reviewer confirms no ideological loading language.

Gate 8: Publication Readiness

  • Evidence quality tier assigned per claim:
    • Tier A strong convergent evidence (multiple high-quality replications)
    • Tier B moderate evidence with known limitations (mixed replication, methodological concerns)
    • Tier C provisional guidance (single study, low-quality evidence, or expert opinion only)
  • Reviewer sign-off completed.
  • Escalation route followed for unresolved disputes (see Escalation and Rejection Policy below).

Red-Flag Reference Table

Use this table during content review. When a red-flag phrase or pattern is detected, replace it with the preferred alternative or add the required qualifier.

Red FlagWhy It Is ProblematicPreferred Alternative
"Studies prove that…""Prove" implies certainty science cannot provide"Research suggests…" / "Evidence indicates…"
"Revolutionary / groundbreaking / game-changing"Hype language; overstates impactDescribe the specific finding without superlatives
"Guaranteed to…"Implies certainty in outcomes"Associated with improvement in…"
"Cures / eliminates / fixes"Causal and absolute; rarely warranted"May reduce…" / "Is linked to lower rates of…"
"All experts agree…"Rarely true; suppresses legitimate disagreement"A broad consensus suggests…" or cite specific bodies
"The science is settled"Appropriate only for a narrow set of topics; misapplied as rhetorical closure"Current evidence strongly supports…"
"As AI confirms…" or citing AI output as a sourceAI output is not a primary sourceCite the underlying study or consensus document directly
"Multiple studies show…" without citationsVague claim with no traceabilityCite specific studies; summarize convergence explicitly
"The only explanation is…"Dismisses alternative interpretations"One well-supported explanation is…"
"This approach transforms / supercharges…"Marketing language; not scientificDescribe the measured outcome and effect size
"Proven brain hack / life hack"Trivializes evidence and implies effortless resultsDescribe the mechanism and the conditions of effectiveness
"100% of participants…" without n and contextAbsolute claim; small-sample artefactReport the actual n and confidence interval
"Never / always" as factual statementsOvergeneralization from limited samplesQualify: "In most cases…", "Evidence suggests…"
"Since time immemorial / ancient wisdom proves…"Appeals to tradition; not evidenceReference specific cultural practice and note evidence status
Uncited paragraph following an AI-generated summaryAI content laundered as authoritativeAdd source citation; manually verify each factual claim

Escalation and Rejection Policy

[HUMAN-REQUIRED] The thresholds and final policy below must be reviewed and approved by a qualified human reviewer before repository-wide adoption. The draft criteria below are provided for human review.

Escalation triggers

Any of the following requires escalation to the Method Reviewer and Approver before publication:

  1. A claim is safety-sensitive (e.g., mental health, medical, child welfare) and rests on Tier C evidence or below.
  2. A claim contradicts established clinical guidelines without a compelling methodological rationale.
  3. Evidence is contested: reviewers disagree on tier classification after one revision cycle.
  4. A conflict of interest in the source cannot be ruled out and the claim is non-trivial.
  5. A single pre-registration or replication failure contradicts a module claim.
  6. AI-generated content cannot be fully traced to primary sources after manual review.
  7. Counter-evidence with comparable methodological quality to the supporting evidence has been identified.

Rejection criteria

A claim must be rejected and may not be published in its current form if:

  1. No credible primary or secondary source can be located after a documented search.
  2. The only supporting sources are retracted, superseded, or conflict-of-interest-compromised.
  3. The claim implies therapy, diagnosis, or clinical intervention without clinical oversight.
  4. The language cannot be revised to accurately reflect the evidence without losing instructional utility.
  5. Safety reviewers have flagged the claim as potentially harmful without a viable mitigation.

Remediation path

OutcomeRequired action
Escalated, resolvableRevise claim language to match evidence tier; re-review within 5 business days
Escalated, unresolvableArchive claim with dissent note; do not publish; open follow-up research task
RejectedRemove claim; document reason in evidence log; notify content author
Safety-flaggedRemove immediately; escalate to Safety Reviewer within 24 hours

Governance integration

This policy integrates with:


Required Artifacts

  • Evidence log (claim → source → source class → quality tier → caveat)
  • Reviewer notes and decision log
  • Final claim language approved for publication
  • Escalation or rejection record (if applicable)