Why Does a Study Become Easier to Believe When Its p-Value Falls Below 0.05?

When Numbers Start Making Decisions · Season Five, “Things That Have Not Happened Yet” · Article 10

1. There is no discovery cliff between 0.049 and 0.051

Imagine two studies with similar sample sizes, methods and estimated effects. One reports p=0.049 and calls the result statistically significant. The other reports p=0.051 and calls it not significant. The first is more likely to enter a headline and subsequent citations; the second may be summarised as showing no effect.

The difference is 0.002. The evidence has not undergone a physical transformation at 0.05. The classification has. A continuous statistic is attached to significant or not significant and then to publication, funding, clinical attention and public trust.

The American Statistical Association’s statement rejects scientific or policy conclusions based only on whether a p-value crosses a threshold. A p-value does not measure the probability that the hypothesis is true, the size of an effect or its importance, and is not by itself a good measure of evidence. American Statistical Association: Statement on p-Values

Why, then, does 0.05 still possess so much authority?

2. What a p-value answers

Given a specified statistical model and null hypothesis, a p-value describes the probability of observing a result at least as extreme as the one obtained. The direction matters: it addresses how incompatible the data are with a model under assumptions. It is not the probability of the hypothesis given the data.

P=0.03 does not mean the null hypothesis has a 3% chance of being true. It does not mean the result has a 97% chance of replication, or that there is only a 3% chance the result arose “by chance”. Those questions require information about design, power, effect size, bias, prior knowledge and other studies.

The same p-value can arise from a huge sample and a negligible effect or from a small sample and a large but imprecise effect. A drug can produce statistical significance with no meaningful clinical benefit. An important effect can fail to cross 0.05 in an underpowered study without being disproved.

The measurement object is narrow: compatibility between data and a specified model. Calling that a line between true and false grants the statistic a power it was not designed to exercise.

3. The value and limit of a common rule

Research communities need rules. If every investigator chooses an evidentiary standard after seeing the outcome, preference will govern. A pre-specified alpha such as 0.05 can control the long-run Type I error rate in a repeated procedure when the null and model assumptions hold.

Binary action is sometimes unavoidable in confirmatory trials, quality control and regulation. A medicine either moves to the next stage or it does not; a study either meets a pre-specified primary endpoint or it does not. An advance rule reduces room for post hoc invention.

The problem is migration. An action boundary created for repeated error control becomes a truth boundary for one study. “Did not meet the pre-specified criterion” becomes “there is no effect”; “met the criterion” becomes “the fact is established”. A procedural threshold gains epistemic authority.

Nor is 0.05 a natural constant for all sciences. Different consequences, fields and multiplicity structures justify different standards. Legitimacy comes from error costs and design, not from the decimal itself.

4. The threshold reshapes research

When publication and careers reward p<0.05, the structure induces behaviour even without deliberate fraud. Researchers may try multiple outcomes, subgroups, time points and models, report only those crossing the line, inspect data repeatedly or redescribe exploratory analyses as pre-specified.

If twenty independent tests each use alpha=0.05 and every null is true, the chance of at least one significant result is much higher than five per cent. Without adjustment, the boundary no longer performs its claimed error control.

Selective publication makes threshold-crossing studies more visible. The literature contains the successes rather than all attempts. Effect sizes can then suffer a winner’s curse: in a small study, only estimates helped enough by random variation cross the line, so published magnitudes tend to be inflated.

A number designed to discipline judgement begins to alter the object being measured. The threshold changes which future knowledge is produced, preserved and rewarded.

5. Comparing studies on either side

The reasonable comparison between p=0.049 and p=0.051 examines estimated effects, uncertainty intervals, sample size, measurement, protocol and bias. Their intervals may overlap substantially and support nearly identical plausible effect ranges.

Confidence intervals are not an automatic solution, but they reveal magnitude and precision. Clinical and policy judgement also needs a minimally important difference, cost, harm and external validity. Statistical significance is not practical importance.

Replication and evidence synthesis matter more than one threshold. Several independent studies with consistent directions and varied methods can collectively support a conclusion even when some p-values exceed 0.05. One spectacular result from exploratory multiple analysis deserves caution even below 0.05.

“Not significant” also does not prove equivalence. Showing two treatments are sufficiently similar requires a pre-specified equivalence or non-inferiority margin and a suitable design. Failure to reject an ordinary null is not such a demonstration.

Another common mistake is to compare a significant result in one subgroup with a non-significant result in another and declare that the subgroups differ. The difference between “significant” and “not significant” is not itself necessarily statistically significant. A proper interaction or direct comparison is required. This matters in clinical and social research, where a threshold can manufacture apparent differences between age groups, sexes or regions. Once such a difference enters a headline, it can guide treatment or policy even though the study never provided adequate evidence for the contrast.

6. Retaining p-values while restricting their power

Confirmatory studies should pre-register primary hypotheses, sample size, exclusions and analysis. Exploratory work remains valuable when clearly labelled rather than rewritten after results appear.

Reports should provide exact p-values with effect estimates and uncertainty intervals, not translate 0.049 as effective and 0.051 as ineffective. Abstracts and media releases should not use “significant” as the conclusion.

Multiplicity, missing data, stopping rules and model assumptions should be addressed, with data and code available for reasonable review. Transparency is a condition of interpretability, not an optional virtue.

Editors and funders should assess questions and methods rather than treat a positive result as entry qualification. Otherwise, individual restraint cannot overcome an incentive system that manufactures bias. High-consequence policy also needs evidence quality, magnitude, reversibility, distribution and alternatives—not one p-value.

7. Scientific facts also form through institutions

A p-value does not emerge automatically from raw data. The research question, null hypothesis, measurement, model, sample and analysis plan constitute it. The 0.05 convention then connects it to journals and policy.

Science needs statistical interfaces because no reviewer can repeat every experiment. Communities distribute trust through methods, standards and replication. When the interface becomes too simple, an editorial system can label p<0.05 “positive” without anyone assessing the evidence as a whole.

Authors and users retain the human commitment to state the strength of their conclusions. “Conventionally significant” cannot become a shield against judgement.

Conclusion: 0.05 can manage a procedure but cannot divide truth

P-values should not be prohibited. With pre-specified hypotheses and a clear error structure, they provide useful information. Pre-set alpha can also be justified in confirmatory procedures requiring binary action.

Scientific writing should stop using 0.05 to separate existence from non-existence. P-values belong beside effect size, intervals, design, prior knowledge and replication. Exploration and confirmation need different labels. Policy must ask more than whether a line was crossed.

There is no discovery cliff between 0.049 and 0.051. What needs to change is the institutional habit that lets a convenient procedure determine which possible future knowledge is worthy of belief.

Primary sources


Discover more from Geoffrey Chen

Subscribe to get the latest posts sent to your email.