Fiche de révision : Understanding Statistical Errors and Reproducibility

Course Outline

  1. Statistical Errors and Hypothesis Testing
  2. Type I and Type II Errors
  3. P Values and False Positives
  4. Reproducibility Crisis in Psychology
  5. Publication Bias and Transparency
  6. Multiple Comparisons and HARKing
  7. Researcher Degrees of Freedom
  8. False Positive Psychology and Study Design
  9. Reporting Guidelines for Authors and Reviewers

1. Statistical Errors and Hypothesis Testing

Key Concepts & Definitions

  • Null hypothesis stance : A null hypothesis stance states a default claim that there is no effect to compare against.
  • Rejection of the null : Rejecting the null means choosing to abandon the no-effect stance based on the test outcome.
  • True positive : A true positive is a decision that matches the presence of the effect when the study finds an effect.
  • True negative : A true negative is a decision that matches the absence of the effect when the study does not find an effect.

Essential Points

  • Frequentist statistical tests evaluate the NULL stance, which claims there is no effect to find.
  • Rejecting the null corresponds to concluding there is an effect rather than accepting the no-effect stance.
  • Correct outcomes occur as true positives when an effect exists and as true negatives when no effect exists.
  • The “cat” logic maps study findings to the four outcomes: true/false positives and true/false negatives.

Memory Hook

Think four boxes: Effect or No effect × Found or Not found.

2. Type I and Type II Errors

Key Concepts & Definitions

  • Type 1 error : A Type 1 error is concluding an effect exists when the true state is that no effect exists.
  • Type 2 error : A Type 2 error is concluding no effect exists when the true state is that an effect exists.
  • Alpha inflation : Alpha inflation is the unintended increase in false-positive findings when decision-making is lenient or repeated.

Essential Points

  • Type 1 error occurs when no cat exists but the study finds a cat, meaning a significant effect despite absence of an underlying effect.
  • Type 2 error occurs when a cat exists but the study finds no cat, meaning an effect exists but is not detected.
  • If the null is true, rejecting it creates a Type 1 error.
  • If the null is false, failing to reject it creates a Type 2 error.

Memory Hook

Type 1: false alarm; Type 2: missed detection.

3. P Values and False Positives

Key Concepts & Definitions

  • P value : A p value is the probability of obtaining results at least as extreme as the observed data under the assumption that the null is true.
  • Statistical significance : Statistical significance is the decision rule that treats sufficiently small p values as evidence against the null.
  • Type 1 error chance : Type 1 error chance is the proportion of times false alarms would occur under the null given the chosen alpha threshold.

Essential Points

  • A p value of 0.05 means there is a 5% chance the observed differences could be noise arising under true absence of the effect.
  • When p < 0.05, the procedure treats the result as unlikely to be a Type 1 error.
  • Interpreting p = 0.05 as “5% chance of a Type 1 error” matches the course’s error framing for the null being true.
  • “Significant at p ≤ 0.05” corresponds to a single-test false-positive probability of about 0.05 under the course’s setup.

Memory Hook

p=0.05 ⇒ think “5% noise under the null.”

4. Reproducibility Crisis in Psychology

Key Concepts & Definitions

  • Reproducibility failure rate : A reproducibility failure rate is the fraction of studies that fail to replicate findings when re-tested by independent efforts.
  • Reproducibility Project : The Reproducibility Project is a collaborative attempt to reproduce prior results and estimate replication failure rates.
  • Social Psychology Special Issue : The Social Psychology Special Issue presents replication findings with an estimated high failure rate.

Essential Points

  • The Reproducibility Project reported an approximate 60% failure rate (Open Science Collaboration, 2015).
  • A Social Psychology Special Issue reported an approximate 70% failure rate (Nosek & Lakens, 2014).
  • Cancer Cell Biology reported an approximate 90% failure rate (Begley & Ellis, 2012).
  • Cardiovascular health reported an approximate 75% failure rate (Prinz, Schlange & Asadullah, 2011).

Memory Hook

Different fields, same pattern: high reported non-replication.

5. Publication Bias and Transparency

Key Concepts & Definitions

  • Publication bias : Publication bias is selective reporting or publication that favors studies with significant-looking outcomes over null results.
  • Lack of transparency : Lack of transparency is missing or incomplete disclosure of methods, analyses, and decisions that affect results.
  • Human error : Human error is unintentional mistake in data handling, analysis, or reporting that can distort conclusions.

Essential Points

  • The “how did we get in this mess” list links publication bias to researcher degrees of freedom, lack of transparency, and human error.
  • Silenced or unreported information can change which results appear in print and thus distort the evidence base.
  • Publication bias interacts with transparency failures because undisclosed decisions can make false positives appear plausible.

Memory Hook

If methods stay hidden and significant results get published, false positives get amplified.

6. Multiple Comparisons and HARKing

Key Concepts & Definitions

  • Multiple comparisons problem : The multiple comparisons problem is the increased chance of false positives when many tests are run on the same data.
  • HARKing : HARKing is hypothesizing after results are known, turning post-hoc findings into supposedly pre-planned hypotheses.
  • Garden of forking paths : The garden of forking paths is the idea that researchers face many analytic choices that lead to many possible test outcomes.

Essential Points

  • With alpha = 0.05 and m comparisons, the chance of at least one false positive is given by 1 − (1 − 0.05)^m in the course example.
  • For m = 16 comparisons, the example false-positive probability is about 0.56.
  • The course example shows the false-positive chance rises with more tests: m=5 gives about 0.774 and m=4 gives about 0.815.
  • Many researchers conduct multiple comparisons “mining” for significance and present results as confirmatory hypothesis testing, which enables HARKing.

Memory Hook

More tests ⇒ more false alarms; HARKing rebrands noise as prediction.

7. Researcher Degrees of Freedom

Key Concepts & Definitions

  • Researcher degrees of freedom : Researcher degrees of freedom is flexibility across research steps that can create multiple equally justifiable choices for hypotheses, data handling, and analysis.
  • Analytic flexibility : Analytic flexibility is the freedom to make multiple analysis choices that can affect which results become statistically significant.

Essential Points

  • Researcher degrees of freedom spans hypothesis generation, study design, data processing, analysis, interpretation, and reporting.
  • Because theories or evidence can be imprecise, multiple decisions may be equally justifiable, which creates room for selective outcomes.
  • The term can be used for opportunistic use of flexibility to achieve desired results via choices such as in- or excluding data.
  • Related terms listed include model uncertainty, multiverse analysis, and P-hacking.

Memory Hook

Degrees of freedom = many valid doors through the same dataset.

8. False Positive Psychology and Study Design

Key Concepts & Definitions

  • False-positive psychology : False-positive psychology studies show how easy it can be to obtain statistically significant evidence for false hypotheses using flexible choices.
  • Computer simulations : Computer simulations are model-based procedures that reproduce the effect of analytic choices and sampling variability on significance rates.
  • Exact replication : Exact replication repeats a study closely enough to test whether results depend on arbitrary analytic decisions.

Essential Points

  • The article states that flexibility in data collection, analysis, and reporting can dramatically raise actual false-positive rates beyond the nominal ≤ 0.05 target.
  • It claims that researchers can be more likely to find evidence for an effect that is false than evidence that no effect exists.
  • Guideline 1 for authors requires choosing a data-collection termination rule before data collection begins and reporting it in the article.
  • Guideline 5 for authors requires reporting statistical results if eliminated observations are included, so effects are not dependent on exclusions.
  • Guideline 6 requires reporting results both with and without a covariate when an analysis includes it.

Memory Hook

Design safeguards aim to stop “significance-by-choice.”

9. Reporting Guidelines for Authors and Reviewers

Key Concepts & Definitions

  • Author reporting requirements : Author reporting requirements are explicit obligations to disclose planned rules, collected variables, conditions, and the impact of analytic choices.
  • Reviewer evaluation standards : Reviewer evaluation standards are criteria for checking compliance, questioning analytic dependence, and deciding on exact replication needs.
  • Data inclusion transparency : Data inclusion transparency means disclosing what happened when observations are excluded or when covariates are removed.

Essential Points

  • Author guideline 2 requires collecting at least 20 observations per cell or giving a compelling cost-of-data-collection justification if fewer are used.
  • Author guideline 3 requires listing all variables collected.
  • Author guideline 4 requires reporting all experimental conditions, including failed manipulations.
  • Reviewer guideline 3 requires authors to demonstrate results do not hinge on arbitrary analytic decisions.
  • Reviewer guideline 4 says that if justifications are not compelling, reviewers should require an exact replication.

Memory Hook

Authors disclose; reviewers probe whether findings survive arbitrary choices.

Key Dates

DateEvent
2015Open Science Collaboration reported an approximate 60% reproducibility failure rate
2014Nosek & Lakens reported an approximate 70% failure rate for a Social Psychology Special Issue
2012Begley & Ellis reported an approximate 90% failure rate for Cancer Cell Biology
2011Prinz, Schlange & Asadullah reported an approximate 75% failure rate for cardiovascular health
1959Karl Popper was cited in relation to the importance of Type 1 error in psychology
2017Chen et al. presented the multiple-comparisons likelihood example
2013Gelman and Loken were referenced in relation to researcher degrees of freedom
2011Simmons et al. were referenced in relation to researcher degrees of freedom
2016Wicherts et al. were referenced in relation to researcher degrees of freedom

Synthesis Tables

Type I vs Type II errors

Error typeNull true?Effect truly exists?
Type 1 errorYesNo; rejecting creates false positive
Type 2 errorNoYes; not rejecting creates false negative

Common Pitfalls & Confusions

  1. Mixing up Type 1 and Type 2 errors is easy because one is a false alarm and the other is a missed detection.
  2. Interpreting p = 0.05 as a 5% chance the null is true confuses the course’s “noise under the null” meaning.
  3. Ignoring multiplicity can make you underestimate false positives when many tests are run on the same dataset.
  4. Treating arbitrary analytic choices as legitimate pre-planned decisions lets researcher degrees of freedom create significance.
  5. Failing to report excluded observations or covariate-removed analyses can hide dependence of results on selective decisions.
  6. Assuming a significant result implies a real underlying effect ignores how flexibility and publication bias can inflate false positives.

Exam Checklist

  1. Explain the NULL stance and how frequentist testing uses it to decide reject vs accept.
  2. Identify which error corresponds to “no effect exists but an effect is found” and which to “effect exists but is not found.”
  3. State how the course frames alpha and what a p value of 0.05 means in terms of Type 1 error chance.
  4. Compute or apply the multiple-comparisons false-positive probability formula 1 − (1 − 0.05)^m for the given m.
  5. Explain what HARKing is and how it relates to presenting mined significance as confirmatory evidence.
  6. Define researcher degrees of freedom as flexibility across research steps and list its main stages from the course.
  7. Recognize examples of related terms such as analytic flexibility and garden of forking paths as linked notions.
  8. Apply the course’s author guideline about pre-specifying the data-collection termination rule before collecting data.
  9. Check whether authors meet the reporting requirements for collected variables, experimental conditions (including failed manipulations), and reporting results with and without covariates.
  10. Use reviewer guidelines to decide when authors must justify analytic decisions or run an exact replication.

Teste tes connaissances

Teste tes connaissances sur Understanding Statistical Errors and Reproducibility avec 11 questions à choix multiples et corrections détaillées.

1. In frequentist hypothesis testing, what does the null hypothesis stance represent?

2. What does a null hypothesis stance represent in statistical hypothesis testing?

Faire le QCM →

Révisez avec les flashcards

Mémorisez les concepts clés de Understanding Statistical Errors and Reproducibility avec 9 flashcards interactives.

Null hypothesis — role?

Default claim of no effect.

Null hypothesis (statistical)

States no effect or difference.

Type I error — definition?

False positive; effect found when none exists.

Voir les flashcards →

Cours similaires

Crée tes propres fiches de révision

Importe ton cours et l'IA génère fiches, QCM et flashcards en 30 secondes.

Générateur de fiches