ScribeLab Writer
Get a Quote

My Sample Size Is Too Small: How to Justify a Constrained Sample and Defend It

Written by Dr. Alina Grace

Published July 27, 2026 · 22 min read

My Sample Size Is Too Small: How to Justify a Constrained Sample and Defend It

You planned for 128 participants and finished with 41. Or your entire population was one clinic, one unit, one cohort, and there was never any prospect of reaching the number your power analysis demanded. Either way, data collection is closed, the number is what it is, and you are staring at a results chapter, wondering whether your whole study is about to be called inadequate. This is one of the most common and most stressful positions a doctoral candidate can be in. The instinctive response, running a post hoc power analysis to show the study was "just underpowered," is the one thing you should not do, because it is a recognized statistical fallacy that examiners spot.

There is a legitimate path through this. A constrained sample is not automatically a fatal flaw. What sinks a study is not a small N but an unjustified one, paired with claims the data cannot support. The defensible response is to reframe the question from "was my study big enough?" to "what can my data actually tell us, and what can it not?" Answer that precisely with a sensitivity analysis and confidence intervals, and report the whole thing transparently. This guide shows you how to do exactly that. When you want the analysis run and the justification written into your chapter, our dissertation data analysis service handles precisely this situation.

One thing should be said plainly at the start. The purpose here is honest justification and correct reporting of a truly constrained sample. None of this can make an inadequate study look adequate, and it should not be used to try. If your data truly cannot address your primary question, the correct conclusion is that the question remains open, which is itself a legitimate contribution.

Quick Answer:

Do not run a post hoc (observed) power analysis. Because observed power is a direct mathematical function of the p-value, a non-significant result always produces low observed power, so it adds no information and cannot excuse a null finding. Instead, run a sensitivity power analysis, which fixes your achieved sample size, alpha, and desired power, and returns the minimum effect size your study could detect. Then, justify the sample using a recognized approach, most often resource constraints or planning for precision, report every result as an effect size with a confidence interval, and write a substantive limitations section stating which inferences your data support and which it does not. Report both your intended and achieved sample sizes, as APA reporting standards require.

Why Post Hoc Power Is the Wrong Move

Start here, because this is the mistake almost everyone reaches for first, and understanding why it fails will save you from it.

A post hoc, or observed, power analysis takes the effect size you actually observed in your completed study and calculates how much power you had to detect it. It feels like it should demonstrate that your non-significant result was a power problem rather than a real absence of effect. It cannot do that, and the reason is mathematical rather than a matter of opinion. Observed power is a one-to-one function of the p-value: as the p-value rises, observed power falls, in lockstep. Knowing the observed power tells you nothing you did not already know from the p-value.

The consequence is fatal for the argument people want to make with it. Any non-significant result, by definition, corresponds to low observed power. So, reporting low observed power after a null finding is circular; you are simply restating that the result was not significant, in different units. Hoenig and Heisey set this out definitively in their well-known critique, showing that computing power after observing the p-value should change nothing about how you interpret that p-value (Hoenig & Heisey, 2001). Levine and Ensom reached the same conclusion in a clinical context, arguing that confidence intervals, not post hoc power, are what properly convey whether an inadequate sample left clinically important effects unexcluded.

This matters practically because committee members and reviewers sometimes ask for post hoc power, not realizing that it is criticized. The professional response is not to refuse but to substitute: provide a sensitivity power analysis and effect sizes with confidence intervals, which answer the underlying question: could this study have detected a meaningful effect, without the fallacy. Being able to explain that distinction calmly is itself a strong signal of statistical maturity, and our dissertation defense preparation guide covers how to handle exactly this kind of challenge.

The Right Tool: Sensitivity Power Analysis

A sensitivity power analysis is the legitimate thing to compute after data collection when your sample size is fixed. Instead of asking how much power you had for the effect you happened to observe, it asks a different and far more useful question: given the sample I actually have, what is the smallest effect I could have detected reliably?

The logic follows from how power analysis works. Four quantities are interlocked: sample size, effect size, alpha, and power, and fixing any three determines the fourth. An a priori analysis fixes effect size, alpha, and power to solve for sample size. A sensitivity analysis fixes sample size, alpha, and power to solve for the effect size, giving you what is called the minimum detectable effect. Crucially, unlike observed power, this does not depend on your observed result at all, so it is not circular. It describes the resolving power of your design.

Running one in GPower takes a minute. Select your test family and the specific statistical test you used, for example, a t-test for a difference between two independent means. Under Type of power analysis, choose "Sensitivity: Compute required effect size, given alpha, power, and sample size." Enter your alpha, usually .05 two-tailed, your desired power, and your actual group sizes. GPower returns the critical effect size your design could detect. It is worth running it at several power levels, .80, .90, and .95, so you can describe your study's sensitivity across a range rather than at a single arbitrary point.

Then interpret it with care, and this is where the real work happens. Compare your minimum detectable effect to the effects reported in comparable studies in your field. If your study could detect effects of the size typically found in your area, your sensitivity was adequate, and the "underpowered" worry is overstated. Suppose your minimum detectable effect is substantially larger than what the literature reports; say so directly. Your study could not reliably detect effects of the size that plausibly exist, so your non-significant result should be read as uninformative about effects that small rather than as evidence against them.

Here is a template you can adapt for your methods or limitations section. Because the sample size was fixed by the available population at the study site, a sensitivity power analysis was conducted in G*Power 3.1. With the achieved sample size, alpha of .05 two-tailed, and power of .80, the minimum detectable effect was d equals the value you obtained. The study was therefore adequately sensitive to effects of that size or larger, but could not reliably detect smaller effects. Results are interpreted with effect sizes and confidence intervals accordingly.

Six Ways to Justify a Sample Size, Only One of Which Is Power Analysis

A widespread misconception among candidates is that an a priori power analysis is the only legitimate way to justify a sample size. It is not. Daniel Lakens sets out six distinct approaches in his comprehensive treatment of sample size justification, and knowing them matters enormously when your N was constrained, because several apply directly to your situation (Lakens, 2022).

The six are: measuring almost the entire population; choosing a sample size based on resource constraints; performing an a priori power analysis; planning for a desired accuracy or precision; using heuristics; and explicitly acknowledging the absence of a justification. Two of these are the workhorses for a constrained sample.

Resource constraints are the most relevant and the most underused. Lakens argues that resource limitations are omnipresent and are always at least a secondary justification for any study's sample size. A responsible researcher evaluates them explicitly rather than pretending they do not exist. Framing your sample this way is not an excuse; it is a recognized justification, provided you do the accompanying work. Lakens is specific about what that work involves. You should address whether the data would still have value in a future meta-analysis, since even a small well-reported dataset contributes to cumulative knowledge. You should report the critical effect size, the smallest effect that would have reached significance given your N. You should report the expected width of your confidence intervals. And you should present a sensitivity analysis across a range of power levels. Do those things, and a resource-constrained sample becomes a defended methodological position rather than an admission.

Planning for accuracy or precision is the second, and it reframes the goal entirely. Instead of asking whether you had enough power to reject a null hypothesis, it asks whether your estimate of the parameter is precise enough to be useful, which is measured by the width of your confidence interval. Maxwell, Kelley, and Rausch made the case that the field over-focuses on power at the expense of accurate parameter estimation, and that estimating an effect precisely is often the more valuable scientific goal (Maxwell, Kelley, & Rausch, 2008). For a completed study with a fixed N, this becomes your interpretive frame: report the estimate with its confidence interval and discuss what that interval includes and excludes.

Measuring almost the entire population deserves a mention because it applies more often than candidates realize. If your study covers nearly everyone in a defined, finite population, one unit's entire nursing staff, every patient meeting rare criteria at your site, the sample size justifies itself, because there is no larger population you are generalizing to. If this describes your study, say so explicitly, since it transforms the framing from a shortfall into a census.

Table 1: Six Ways to Justify a Sample Size (Lakens, 2022)

Justification

When It Applies

Strength for a Constrained Sample

Almost the entire population

You measured nearly everyone in a finite population

Very strong; the sample justifies itself

Resource constraints

Time, access, funding, or setting limited recruitment

Strong, if paired with sensitivity analysis and CI width

A priori power analysis

Planned before data collection

Not available once collection has closed

Planning for accuracy / precision

The question is about the size of a parameter

Strong as an interpretive frame; report CI width

Heuristics

Rules of thumb (for example, 20 per cell)

Weak; rarely persuades examiners

No justification, acknowledged

No principled basis existed

Weakest, but better than inventing one

Reporting Non-Significant Results Correctly

If your constrained study produced a non-significant result, how you write about it determines whether your examiner sees a competent researcher or an overclaiming one. There are three correct moves and one common error.

The error is concluding that there is no effect. A non-significant result means you did not detect an effect, not that no effect exists. Altman and Bland put this into a phrase that has become standard: absence of evidence is not evidence of absence (Altman & Bland, 1995). Writing that your intervention "had no effect" when your test was merely non-significant is an inferential error examiners catch immediately. It is more damaging than the small sample itself.

The first correct move is to report the effect size with its confidence interval and interpret the interval substantively. This is far more informative than a p-value, because it shows both your best estimate and the range of values compatible with your data. A wide interval that spans both trivial and substantial effects tells your reader plainly that the study was not precise enough to distinguish them. An interval that excludes large effects, even while including zero, tells them something real: whatever is happening, it is probably not large.

The second correct move, where your precision allows it, is equivalence testing. Standard hypothesis testing can never confirm a null; it can only fail to reject it. Equivalence testing, typically via the two one-sided tests procedure, lets you test whether an effect is smaller than the smallest effect size of interest that you define in advance. That allows you to positively conclude any effect is negligible, rather than merely failing to find one (Lakens, Scheel, & Isager, 2018). This is the rigorous way to argue for an absence. One honest caveat: equivalence testing still requires enough precision relative to your smallest effect of interest. With a very small sample, you may be able to reject neither the null nor the equivalence bound, in which case the correct conclusion is that the result is inconclusive, and saying so is better than forcing a claim.

The third move is framing. Report your finding as what it is, an inconclusive or imprecise result on this question, and note the contribution your data still makes, including to future meta-analyses. That is a defensible scientific position; "we found no difference" from an underpowered test is not.

Ended up with a smaller sample than you planned?

Send us your data and your original power analysis. A PhD statistician will run the sensitivity analysis, recompute your results as effect sizes with confidence intervals, and draft the sample-size justification and limitations section that your committee will accept, with every claim matched to what your data support. Request a sample size review and receive an itemized quote within 2 to 4 business hours, no obligation.

If You Got a Significant Result From a Small Sample, Be Careful

This section is counterintuitive, and it is the one most likely to impress an examiner, because almost no candidate raises it voluntarily. If your small study produced a statistically significant result, that is not straightforwardly good news, and treating it as vindication is a mistake.

Low statistical power not only increases false negatives. It also degrades the significant findings you do get, in two specific ways. First, it lowers the probability that a significant result reflects a true effect, the positive predictive value of your finding. Second, and more concretely, it inflates the apparent size of effects that reach significance. The mechanism is simple once seen: when power is low, and an effect must clear a significance threshold to be detected, the only effects that clear it are the ones that happen, by chance, to be overestimated. This is often called the winner's curse. Button and colleagues documented this in their influential analysis of power in neuroscience. They estimated the median power of studies in the field at around 21 percent, and showed that effect estimates from studies powered in that range are likely to be substantially inflated (Button et al., 2013).

The practical implication for your write-up is direct. If you report a significant result from a small sample, present it with its confidence interval, note that the point estimate is likely an overestimate of the true effect, and avoid building strong practical recommendations on its magnitude. A candidate who raises this unprompted demonstrates real understanding of their own limitations, which is exactly what a viva is testing.

Legitimate Design and Analytic Responses

Beyond justification and reporting, there are genuine methodological moves that improve what a small sample can deliver. Some are available only if you can still adjust your design; others apply at the analysis stage.

Within-subjects and repeated-measures designs are the single most powerful lever, because measuring the same participants under multiple conditions removes between-person variance from the comparison. A study with few participants but many observations each can have considerably more power than a between-subjects design with far more people supplying one observation apiece. Reducing measurement error is a second and often overlooked lever: unreliable measures attenuate effects and cost power, so improving the reliability of your instrument can be equivalent to gaining participants. Including covariates through analysis of covariance reduces error variance and increases precision without recruiting anyone. Bootstrapping produces confidence intervals without relying on large-sample assumptions, which suits small and irregular samples. Bayesian estimation offers an alternative framework that does not depend on fixed-N power and can incorporate prior information, though it requires priors you can defend. Choosing among these responses depends on your design and how much of it remains adjustable, which is one of the judgments our dissertation support helps candidates make. And supplementing with qualitative data is a recognized strategy for inherently small or rare populations, adding depth where breadth is impossible.

For truly small or rare populations, single-case and small-N experimental designs deserve serious consideration, because they are rigorous designs with formal standards rather than compromises. They are adaptations of interrupted time-series logic and can provide credible experimental evaluation of intervention effects in populations too small, too heterogeneous, or too atypical to form conventional groups. If your population is inherently tiny, choosing a design built for that situation is stronger than running an underpowered group comparison.

One method worth flagging with caution: switching to a one-tailed test increases power for a directional hypothesis, but it is controversial, because it can be used to manufacture significance, and it forfeits the ability to detect effects in the unexpected direction. It is defensible only when specified in advance with a strong theoretical rationale, never chosen after seeing the data.

What Your Committee Actually Expects

Examiners are more forgiving of a small sample than candidates fear, and far less forgiving of how it is handled. Four things are expected.

Report both your intended and achieved sample size, and explain the discrepancy. This is not optional under APA reporting standards, which require the intended sample size, the achieved sample size where different, and how the sample size was determined (American Psychological Association, JARS-Quant). Stating plainly that you planned for a certain N, achieved another, and why, is more credible than quietly reporting only what you got.

Replace post hoc power with a sensitivity analysis, as above, and report effect sizes with confidence intervals throughout.

Write a substantive limitations section, not a boilerplate sentence. "A limitation of this study is the small sample size" is the version that irritates examiners, because it acknowledges the issue without engaging with it. The strong version specifies which inferences your data support and which they do not. It states the minimum effect you could have detected, notes what your confidence intervals exclude, and flags the risk of effect-size inflation for any significant findings. That paragraph turns a weakness into evidence of judgment, and it is one of the sections our dissertation chapter supports most often.

Finally, do not overclaim anywhere else in the document. The most common way a small-sample study fails a viva is not the sample; it is a discussion chapter and abstract that make population-level claims the data cannot bear. Align every claim with your evidence, a discipline our overview of the ten common dissertation mistakes repeatedly returns to.

Table 2: Responding to a Small Sample — What Works and What Fails

Situation

What Fails

What Works

Explaining a non-significant result

Post hoc / observed power (circular)

Sensitivity analysis plus effect size with CI

Describing what you found

"There was no effect"

"No effect was detected"; interpret the interval

Arguing an effect is negligible

Treating p > .05 as proof of no effect

Equivalence testing (TOST) against a defined SESOI

A significant result from a small N

Treating it as confirmation

Report the CI; note likely effect-size inflation

Writing the limitations section

One boilerplate sentence

Minimum detectable effect, CI exclusions, supported claims

A Note for DNP and Nursing Candidates

If you are running a DNP quality improvement project on a single unit, the framing changes in a way that resolves much of this anxiety. Your project is not a hypothesis test about a population parameter; it is an evaluation of whether a change produced a sustained, non-random shift in a process. That question is answered with time-ordered data, not with power analysis.

Run charts and statistical process control charts are the appropriate tools, and they work with modest numbers of sequential data points rather than large samples. Established guidance on run charts in healthcare recommends collecting around 20 to 30 data points where possible. It also notes that a chart remains useful with far fewer, and that a shift can be signaled with as few as six consecutive points on one side of the median (Perla, Provost, & Murray, 2011). Statistical process control is likewise an established alternative for testing intervention effects using data collected over time, distinguishing common-cause from special-cause variation, which is particularly valuable where randomization and large samples are impossible.

The reporting standard for this work is SQUIRE 2.0, which expects you to describe how you assessed whether outcomes were due to the intervention and how you examined variation over time. Frame your project that way, and the small-N question largely dissolves, because you are justifying a number of data points over time rather than a sample size. Where your project does include a comparative research component, the guidance in the rest of this article applies to that component. Nursing work also puts particular weight on the distinction between statistical and clinical significance: a change can matter clinically without reaching significance in a small unit, and reporting the magnitude in natural units is what communicates that. Our DNP capstone support is built around this distinction.

Frequently Asked Questions

My sample size is smaller than the power analysis required. Is my dissertation ruined?

Almost certainly not. What examiners penalize is an unjustified sample paired with claims the data cannot support, not a small N in itself. The path forward has four steps. Run a sensitivity power analysis to establish the smallest effect your study could detect. Justify the sample using a recognized approach, such as resource constraints. Report every result as an effect size with a confidence interval. Then write a substantive limitations section stating exactly which inferences your data do and do not support.

Should I run a post hoc power analysis to explain my non-significant result?

No. Observed power is a direct mathematical function of the p-value, so a non-significant result always yields low observed power. Reporting it adds no information beyond the p-value and cannot show your study was "just underpowered"; the argument is circular. Hoenig and Heisey established this decisively. If a reviewer asks for post hoc power, substitute a sensitivity power analysis and effect sizes with confidence intervals, which answer the real question without the fallacy.

What is a sensitivity power analysis, and how do I run one?

A sensitivity analysis fixes your achieved sample size, alpha, and desired power, and solves for the minimum effect size your study could reliably detect. In G*Power, select your test, choose "Sensitivity: Compute required effect size," and enter alpha, power, and your actual sample sizes. Unlike post hoc power, it does not depend on your observed result, so it is not circular. Run it at several power levels and compare the result to effects reported in comparable studies.

How do I justify a small sample size in my methodology?

Use a recognized justification rather than an apology. For a truly constrained sample, the resource-constraints justification is usually correct. State plainly that the available population or setting limited the sample, then do the supporting work: report the smallest effect that would have been significant, report your expected confidence interval width, present a sensitivity analysis, and note the data's value for future meta-analysis. If you sampled nearly an entire finite population, say so, since that justifies itself.

How should I write the limitations section about my sample size?

Avoid the boilerplate sentence. A strong limitations paragraph states the minimum effect your study could detect, specifies which inferences the data support and which they do not, notes what your confidence intervals exclude, flags that any significant effect may be inflated, and explains why the sample was constrained. Report both the intended and achieved sample size and explain the discrepancy. This engagement demonstrates judgment; a single sentence acknowledging a small sample without analysis does not.

Can I say there was no effect if my result was not significant?

No. A non-significant result means you did not detect an effect, not that none exists, a principle Altman and Bland summarized as absence of evidence not being evidence of absence. Instead, report the effect size with its confidence interval and interpret what the interval includes and excludes. If your precision allows and you can define the smallest effect of interest, equivalence testing lets you positively conclude that an effect is negligible. If neither can be established, report the result as inconclusive.

I got a significant result with a small sample. Is that a problem?

It requires caution rather than celebration. Low power reduces the probability that a significant result reflects a true effect, and it inflates the size of effects that do reach significance, because only overestimated effects clear the threshold, a pattern known as the winner's curse. Report the finding with its confidence interval, note that the point estimate is likely an overestimate, and avoid building strong practical recommendations on its magnitude. Raising this yourself signals real methodological understanding.

Defending What Your Data Can Actually Support

A constrained sample becomes a serious problem only when it is left unjustified or paired with claims it cannot carry. Run the sensitivity analysis so you know the smallest effect your study could detect. Justify the sample with a recognized approach rather than an apology. Report effect sizes with confidence intervals instead of leaning on significance. Say what your data cannot tell us as clearly as what it can. And never reach for post hoc power to rescue a null. Do that, and you walk into your defense able to state precisely what your study establishes and where its limits lie, which is a stronger position than many candidates with larger samples ever reach.

If you would rather a PhD statistician run the sensitivity analysis, recompute your results with confidence intervals, and write the justification and limitations your committee expects, send us your data and your original power analysis. You will have an itemized quote within 2 to 4 business hours, with no obligation.

About the author

Dr. Alina Grace

Dr. Alina Grace

Meta-Analysis & Synthesis Lead

PhD Epidemiology; MSc Evidence-Based Healthcare

Evidence synthesis lead specializing in PROSPERO-registered systematic reviews and meta-analysis.

View full profile

Ready to Get Your Quote?

Describe your project and a PhD specialist will reply with an itemized quote within 2-4 business hours. No signup, no payment, no obligation.

Prefer email? Send your project details to info@scribelabwriter.com

Chat with us on WhatsApp