ScribeLab Writer
Get a Quote

My Results Were Not Significant: How to Interpret, Write Up, and Defend Null Findings

Written by Dr. Alina Grace

Published July 27, 2026 · 19 min read

My Results Were Not Significant: How to Interpret, Write Up, and Defend Null Findings

You ran the analysis you planned, on the data you worked years to collect, and the p-value came back above .05. Your hypothesis was not supported. If your first thought was that the whole project has failed, you are in good company, and you are also wrong. A doctoral thesis is assessed on the quality of the question, the soundness of the design, and the rigor of the reasoning, not on whether the result crossed a threshold. Null findings are published, defended, and awarded every year.

What sinks a candidate is mishandling the null. That means claiming the intervention "had no effect" when the study merely failed to detect one, reaching for post hoc power to explain the result away, or quietly rewriting the hypothesis. Hence, the data appear to confirm something. Those are the errors examiners catch, and each one is avoidable. This guide shows you how to interpret a non-significant result correctly, how to distinguish a truly informative null from an inconclusive one, what to write in your discussion chapter, and how to defend it. When you want the analysis rechecked and the chapter written to that standard, our dissertation data analysis service does exactly this work.

One boundary, stated up front. Everything here is about interpreting and reporting a null result accurately. None of it is about making a disappointing study look successful. The honest version is also the more defensible one, and examiners can tell the difference.

Quick Answer:

A non-significant result means your data provides little evidence against the null hypothesis. It does not mean the null is true, and writing that your intervention "had no effect" is an inferential error. First, report the effect size with its confidence interval and ask what the interval excludes. If it still includes effects you would care about, your result is inconclusive; if it excludes them, you have an informative null. To claim positively that any effect is too small to matter, use equivalence testing against a pre-defined smallest effect of interest, or a Bayes factor that quantifies evidence for the null. Never use post hoc observed power to explain the result, never rewrite your hypothesis after seeing the data, and report every pre-specified outcome regardless of direction.

The Mistake Almost Everyone Makes

Start with the interpretation itself, because the most common error happens in the first sentence a candidate writes about their result.

A p-value above your alpha threshold tells you that the data are reasonably compatible with the null hypothesis. It does not tell you the null is true. The distinction sounds academic until you see how often it is missed. When researchers examined 137 non-significant findings reported in the 2015 abstracts of three well-regarded psychology journals, they found that in 72 percent of cases, the results were misinterpreted, with authors inferring that the effect was absent. Related work has found that roughly 60 percent of researchers will accept the null on the basis of a non-significant result. These are published researchers in peer-reviewed journals, which tells you how easy the error is to make.

The correct framing was given its enduring phrase by Altman and Bland: absence of evidence is not evidence of absence (Altman & Bland, 1995). Your study did not detect an effect. Whether that is because there is no effect, or because your study could not have detected the effect that exists, is a separate question, and answering it is the intellectual work of your discussion chapter.

This connects to a wider point about what p-values can carry. The American Statistical Association's statement on p-values cautions that scientific conclusions should not be based only on whether a p-value passes a specific threshold, warning that bright-line rules such as .05 can lead to erroneous beliefs and poor decisions (Wasserstein & Lazar, 2016). There is an active and contested debate about how far to push this, with some methodologists arguing for retiring significance thresholds altogether and others defending pre-specified testing. You do not need to take a side in your thesis. You do need to interpret your own p-value correctly.

Why Post Hoc Power Is Not the Answer

The instinct, once a result comes back null, is to compute observed power and argue the study was underpowered. Do not do this. Observed power is a direct mathematical function of the p-value, so a non-significant result always produces low observed power. Reporting it restates the p-value in different units and adds nothing, which is why Hoenig and Heisey describe the practice as fundamentally flawed and label the resulting confusion the power approach paradox (Hoenig & Heisey, 2001).

If your concern is whether the study could have detected a meaningful effect, that is a legitimate question with a legitimate answer: a sensitivity analysis, which reports the minimum effect your design could detect, together with confidence intervals. We cover that machinery in detail in our guide on what to do when your sample size is too small, and the planning-stage version in our G*Power power analysis guide.

Uninformative Null or Informative Null? The Question That Shapes Everything

Before you write a word of discussion, establish which kind of null you have, because the two lead to completely different chapters.

Report your effect size with its confidence interval, then ask a single question: Does the interval exclude the effects you would consider meaningful? A wide interval that spans everything from trivial to substantial effects means your study was not precise enough to distinguish them. That is an uninformative null, and the honest conclusion is that the question remains open. A narrow interval clustered around zero, which excludes effects of practical importance, means something more. That is an informative null, and you can say, with evidence, that any effect is probably too small to matter.

This shift from significance testing toward estimation is what Cumming calls the new statistics: report the effect and its interval, and interpret the range of values compatible with your data rather than a binary verdict. It is also what turns a null result from an apology into a finding. Getting the reporting format right matters here, and our guide on writing the results chapter in APA 7 covers exactly how these numbers should appear.

Table 1: Is Your Null Informative or Inconclusive?

Feature

Uninformative (Inconclusive) Null

Informative Null

Confidence interval

Wide; still includes effects that would matter

Narrow; excludes effects that would matter

Equivalence test (TOST)

Cannot reject the smallest effect of interest

Rejects the smallest effect of interest

Bayes factor

Near 1 (anecdotal; data are insensitive)

Favors the null (conventionally above 3)

What you can claim

The question remains open

Any effect is likely too small to matter

What to recommend

A larger or better-measured replication

Redirecting effort away from this intervention

Making a Positive Claim: Equivalence Testing and Bayes Factors

Standard hypothesis testing has a structural limitation that is worth stating plainly to your committee: it can never confirm a null. It can only fail to reject one. If you want to claim positively that an effect is negligible, you need different tools, and there are two.

Equivalence testing is the frequentist route. Using the two one-sided tests procedure, you specify in advance a smallest effect size of interest, the smallest effect that would matter practically or theoretically, and then test whether your observed effect is significantly smaller than that bound. If it is, you can conclude that any true effect is too small to be meaningful, which is a positive finding rather than a failure to find (Lakens, Scheel, & Isager, 2018). The critical requirement is that the smallest effect of interest must be justified, ideally through an anchor-based method or a minimal clinically important difference in health research, not pulled from a generic benchmark table. In clinical research, this logic is formalized in non-inferiority and equivalence trial designs, which have their own CONSORT reporting extension.

One honest limitation: equivalence testing needs precision. With a small sample, you may be unable to reject either the null or the equivalence bound, in which case the correct conclusion is that the result is inconclusive. Saying so is better than forcing a claim your data cannot support.

Bayes factors are the Bayesian route, and they do something frequentist tests cannot: they quantify evidence for the null relative to the alternative. Dienes makes the case directly, arguing that a Bayes factor lets you distinguish evidence for the null from mere insensitivity, which is precisely the distinction a candidate with a null result needs (Dienes, 2014). Conventional interpretation treats a Bayes factor between 1 and 3 as anecdotal evidence, 3 to 10 as moderate, and above 10 as strong, so a Bayes factor for the null above 3 is commonly read as moderate evidence that the null is favored. Two caveats belong in your write-up. Those thresholds are conventions, not rules, and the resulting value depends on your choice of prior, so you should report a sensitivity analysis showing how the conclusion holds across reasonable priors. The wider Bayesian versus frequentist question remains contested among methodologists, so present your approach as a considered choice, not the only correct one.

Not sure whether your null result is inconclusive or meaningful?

Send us your data and your hypotheses. A PhD statistician will recompute your results as effect sizes with confidence intervals, run equivalence tests or Bayes factors where they apply, and draft a discussion chapter that says exactly what your findings support and nothing more. Request a null results review and receive an itemized quote within 2 to 4 business hours, no obligation.

Explanations Worth Considering in Your Discussion

A strong discussion chapter works through the plausible reasons your study did not detect an effect, and does so with evidence rather than speculation. Consider each of the following against your own design, and discuss the ones that actually apply.

A true null. The effect may not exist, or may be too small to matter. This is a legitimate scientific finding and should be stated first if your evidence supports it.

Insufficient precision. Your sample may have been too small to detect the effect that exists. Support this with your sensitivity analysis and confidence interval width, never with observed power.

Measurement unreliability. This one is underused and impresses examiners. Unreliable measures attenuate observed effects toward zero: under classical test theory, the observed correlation is the true correlation multiplied by the reliabilities of the measures. Hence, a noisy instrument mechanically shrinks what you can detect. If your scale's reliability was modest, this is a substantive explanation, not an excuse. Use it to explain attenuation, not to inflate a reported effect size.

Restriction of range. If your sample varied little on the predictor or outcome, there was less covariation available to detect.

Ceiling and floor effects. If most participants scored near the top or bottom of your measure, real differences had nowhere to appear.

Implementation fidelity or dose failure. In intervention studies, this is often the real story. If the intervention was not delivered as intended, a null result tells you about implementation rather than efficacy, which is a different conclusion entirely. The implementation science literature calls this a Type III error, and the standard framework is explicit that, unless fidelity is evaluated, you cannot tell whether a lack of impact reflects poor implementation or a truly ineffective program (Carroll et al., 2007). If you collected any fidelity data, this belongs in your discussion.

Insufficient follow-up. The outcome may take longer to change than your measurement window allowed.

Heterogeneous effects. An average null can conceal offsetting effects in different subgroups, though any subgroup analysis must be labeled exploratory unless pre-specified.

Confounding. An uncontrolled confounder may have masked a real effect.

Table 2: Explanations for a Null Result, and the Evidence That Supports Each

Explanation

What to Check

Evidence to Cite in Your Discussion

A true null effect

Does the CI exclude meaningful effects?

Equivalence test or Bayes factor result

Insufficient precision

Achieved N and CI width

Sensitivity analysis (never post hoc power)

Measurement unreliability

Reliability coefficients of your measures

Attenuation of effects under classical test theory

Restriction of range

Variance and spread of key variables

Descriptive statistics showing limited variation

Ceiling or floor effects

Distribution of scores near scale limits

Histograms and percentage at scale extremes

Implementation fidelity failure

Was the intervention delivered as intended?

Fidelity data; Type III error framework

Insufficient follow-up

Time needed for the outcome to change

Comparable studies' follow-up periods

Confounding or heterogeneity

Uncontrolled variables; subgroup patterns

Pre-specified analyses only; label others exploratory

Two disciplines apply throughout. Only raise explanations plausible for your design, since a scattergun list reads as defensive. And check that your analysis itself was sound before attributing the null to the phenomenon. That means confirming you chose an appropriate test and that its assumptions hold, covered in our guide on which statistical test to use, and checking and reporting statistical assumptions.

The Temptation You Must Resist

There is a specific and serious integrity risk that arises with null results, and naming it protects you.

When the predicted effect does not appear, it is tempting to look through the data for something that did reach significance, and then present that as though it had been the hypothesis all along. This practice has a name: HARKing, or hypothesizing after the results are known. Kerr's original analysis explains why it is corrosive. When HARKing follows a chance finding, theory gets constructed to explain an effect that is not real, translating Type I errors into the literature (Kerr, 1998). The related temptation is to keep adjusting analytic choices until something crosses the threshold, which the false-positive psychology literature showed can manufacture significance from noise through undisclosed flexibility alone.

The remedy is simple and costs you nothing. Report your pre-specified analysis as confirmatory, exactly as planned, including its null result. Then report anything you noticed afterward as explicitly exploratory, framed as a hypothesis for future work. Generating new hypotheses from data is legitimate science; disguising them as predictions is not. Examiners read for this distinction, and a chapter that separates confirmatory from exploratory analysis cleanly signals integrity rather than weakness. Overclaiming of this kind is one of the recurring themes in our overview of the ten common dissertation mistakes.

Your Null Result Has Real Scientific Value

This is not consolation; it is a documented structural problem that your finding helps correct, and it belongs in your discussion.

Null results are systematically missing from the published literature. Rosenthal named this the file drawer problem, illustrating it with the image of journals filled with the significant studies while the null ones sit in filing cabinets (Rosenthal, 1979). That framing was a rhetorical illustration rather than a measurement, but subsequent work has quantified the problem precisely. Examining a set of peer-reviewed studies with known outcomes, Franco and colleagues found that only about 20 percent of those with null results appeared in print, against roughly 60 percent of those with strong results. More revealingly, around 65 percent of null studies were never even written up (Franco, Malhotra, & Simonovits, 2014). The bias operates mostly at the point where researchers decide not to bother.

The consequence is a literature that overstates how often interventions work. A striking illustration comes from Registered Reports, where methods are peer reviewed and accepted before data collection. When researchers compared these with conventionally published psychology studies, 96 percent of standard reports supported their first hypothesis, against only 44 percent of Registered Reports. The gap is a measure of how much the conventional record is filtered.

Write your null up, and you are correcting that record rather than adding to the distortion. Your data can also contribute to future meta-analyses, where a well-reported null carries real weight. In clinical and nursing fields, there is an ethical dimension too: suppressing findings that an intervention did not work allows ineffective practice to continue looking effective.

What Examiners Actually Expect

Four expectations, and none of them has a significant p-value.

Complete reporting, regardless of direction. The APA reporting standards for quantitative research require you to report all relevant results, whether or not your hypotheses were supported, including findings that run counter to expectations, and to distinguish primary, secondary, and exploratory analyses (American Psychological Association, JARS-Quant). Selective reporting is a serious problem; a null you report fully is not.

Correct inferential language. Say that no effect was detected, not that there was no effect, and reserve claims of equivalence for cases where you tested for it.

A discussion that reasons rather than apologizes. Work through the plausible explanations, say which your evidence supports, and state what would resolve the question.

Consistency across the whole thesis. The most common way a null study fails is not the null itself, but an abstract or conclusion that quietly overclaims. Every claim must match the evidence, which is the alignment our dissertation chapter support and broader dissertation support are built to enforce.

On the question everyone asks: yes, candidates defend and pass with null findings. Examiners assess the research process. Being able to explain why your study was well designed, what it could and could not detect, and what your results contribute is the substance of that defense. Our dissertation defense preparation guide rehearses those questions. Institutional regulations vary, so confirm your own program's requirements.

A Note for DNP and Nursing Candidates

If your DNP quality improvement project did not produce a signal, the framing differs in a way that helps you. A run chart or control chart showing only common-cause variation, no shift, no trend, no point beyond the limits, means no special-cause change was detected. That is a legitimate, reportable result. Charts show whether a process changed, not why, so the absence of a signal should send you toward context and fidelity rather than toward spin. The mechanics are covered in our guide on DNP data analysis with run charts and statistical process control.

SQUIRE 2.0, the reporting standard for improvement work, is built for exactly this situation. It asks you to report what was learned, how context influenced the work, the reasons for any differences between observed and anticipated outcomes, and factors that may have limited internal validity, such as confounding, bias, or imprecision (Ogrinc et al., 2016). A DNP project that documents why an intervention did not take hold in a particular unit is valuable to the next team that tries it.

Nursing also puts particular weight on the distinction between statistical and clinical significance. A change can be clinically meaningful without reaching statistical significance, especially on a single unit, and reporting the magnitude in natural units communicates that far better than a p-value. This is central to our DNP capstone support.

Frequently Asked Questions

My results were not significant. Does that mean my dissertation has failed?

No. A doctoral thesis is assessed on the quality of the research question, the soundness of the design, and the rigor of your reasoning, not on whether a p-value crossed a threshold. Candidates defend and pass with null findings regularly. What causes problems is mishandling the result: claiming there was no effect when you merely failed to detect one, using post hoc power to explain it, or rewriting hypotheses after the fact.

Can I say my intervention had no effect if the result was not significant?

No. A non-significant result means you did not detect an effect, not that no effect exists; the principle Altman and Bland summarized as absence of evidence not being evidence of absence. Research shows this error is very common, with one review finding that 72 percent of non-significant findings in a sample of psychology abstracts were misinterpreted as showing an absent effect. Say that no significant effect was detected, and interpret your confidence interval.

How do I know if my null result is meaningful or just inconclusive?

Look at the confidence interval around your effect size and ask what it excludes. If the interval still includes effects large enough to matter practically, your study was not precise enough to rule them out, and the result is inconclusive. If the interval excludes meaningful effects and clusters near zero, you have an informative null and can say any true effect is probably too small to be important.

How can I show there is truly no effect?

Standard hypothesis testing cannot confirm a null; it can only fail to reject it. To make a positive claim, use equivalence testing via the two one-sided tests procedure against a smallest effect size of interest that you justify in advance, or compute a Bayes factor, which quantifies evidence for the null relative to the alternative. Both require adequate precision. If neither can be established, report the result as inconclusive rather than forcing a claim.

What should I write in my discussion chapter about a null result?

State plainly that the hypothesis was not supported, report the effect size and confidence interval, and interpret what the interval excludes. Then work through the explanations plausible for your design: a true null, insufficient precision, measurement unreliability, restriction of range, ceiling or floor effects, implementation fidelity problems, insufficient follow-up, heterogeneous effects, or confounding. Say which your evidence supports and what would resolve the question. Do not use post hoc power.

Is it acceptable to change my hypothesis to match what I found?

No. Presenting a post hoc hypothesis as though it had been predicted is called HARKing, and Kerr's analysis explains why it damages the literature: theory gets built to explain effects that may not be real. Report your pre-specified analysis as confirmatory, including its null result, and report anything you noticed afterward as explicitly exploratory and a question for future research. Generating hypotheses from data is legitimate; disguising them as predictions is not.

Why do null results matter if journals rarely publish them?

Because their absence distorts the evidence base, Franco and colleagues found that only about 20 percent of null studies were published against roughly 60 percent of strong results, and about 65 percent of nulls were never written up at all. The distortion is measurable: 96 percent of conventionally published psychology studies supported their first hypothesis, against 44 percent of Registered Reports, where acceptance precedes results. Reporting your null corrects the record and contributes to future meta-analyses.

Reporting What You Actually Found

A null result is not the end of a doctorate; it is a finding that has to be handled with more care than a significant one. Report the effect size and interval rather than leaning on the p-value. Establish whether your null is informative or inconclusive, and say which. Use equivalence testing or a Bayes factor if you want to claim an effect is negligible. Work through the explanations your design makes plausible, keep confirmatory and exploratory analyses separate, and never let the abstract claim more than the data support. Do that, and you walk into your defense with something many candidates with significant results lack: a precise account of what your study establishes and where its limits lie.

If you would rather a PhD statistician recheck the analysis, run the equivalence or Bayesian tests where they apply, and draft a discussion chapter that claims exactly what your data support, send us your data and hypotheses. You will have an itemized quote within 2 to 4 business hours, with no obligation.

About the author

Dr. Alina Grace

Dr. Alina Grace

Meta-Analysis & Synthesis Lead

PhD Epidemiology; MSc Evidence-Based Healthcare

Evidence synthesis lead specializing in PROSPERO-registered systematic reviews and meta-analysis.

View full profile

Ready to Get Your Quote?

Describe your project and a PhD specialist will reply with an itemized quote within 2-4 business hours. No signup, no payment, no obligation.

Prefer email? Send your project details to info@scribelabwriter.com

Chat with us on WhatsApp