A meta-analysis of observational studies gets criticized for a reason that surprises many candidates: it looks too good. Pool ten cohort studies, the confidence interval around the summary estimate narrows, the result reaches significance, and the finding reads as solid. But if those ten studies share the same confounding, the same selection bias, the same measurement problem, then meta-analysis has not removed the bias. It has averaged it and wrapped it in false precision. Reviewers who work with observational evidence know this, and it is the first thing they check.
Synthesizing observational studies is not a smaller version of synthesizing trials. In a randomized trial, randomization is expected to balance confounders, so pooling mainly averages away random error. In cohort, case-control, and cross-sectional studies, confounding and bias are built into every primary study, and no amount of pooling removes them. That difference changes which reporting guideline you use, which risk-of-bias tool applies, how you handle the estimates, and what you may claim. Get those right, and an observational meta-analysis is as defensible as any trial synthesis. Get them wrong, most often by treating observational evidence like trial evidence, and it comes back.
This guide covers the four things reviewers check: the right reporting standard (MOOSE), the right risk-of-bias tool (ROBINS-E, not the Newcastle-Ottawa Scale used uncritically), the statistical handling that observational data demand, and how to rate this evidence under GRADE. When you want a statistician to build or check that synthesis, our meta-analysis service does exactly this work.
Quick Answer:
If you pool cohort, case-control, or cross-sectional studies, report against MOOSE plus PRISMA 2020, and assess risk of bias with ROBINS-E, not the Newcastle-Ottawa Scale used as a summary score. Pool adjusted estimates, never crude ones, and tabulate which confounders each study adjusted for. Use a random-effects model, report tau-squared, I-squared, and a prediction interval, and expect I-squared to be high. Never claim causation from a significant pooled association that may reflect shared confounding. Under GRADE, observational evidence starts at low certainty but can be rated up for a large effect, a dose-response gradient, or confounding that would work against the observed effect. The three things that get these reviews rejected are no structured risk-of-bias assessment, pooling incompatible estimates, and over-claiming causation.
Why Observational Meta-Analysis Is Harder
Three biases sit inside every observational study, and meta-analysis cannot dissolve any of them. Confounding, where a third factor tied to both exposure and outcome distorts the association, does not shrink as studies are pooled. Selection bias, from how participants entered the study or the analysis, persists. Information bias, from misclassified exposures or outcomes, carries through. Because these biases differ in size and direction across studies, between-study heterogeneity is typically far larger than in a trial synthesis.
The classic warning is Egger and colleagues' "spurious precision," which showed that pooling observational data can generate a narrow, confident-looking confidence interval around a point that is biased, because the component studies share systematic error (Egger, Schneider, & Davey Smith, 1998). This is the garbage-in, garbage-out problem, and it has a direct consequence for what you may write: a statistically significant pooled estimate summarizes an association, not a causal effect. Whether that association is causal is a separate question that the meta-analysis alone cannot answer, and saying so explicitly in your discussion is one of the things reviewers look for. The general mechanics of pooling are covered in our guide on how to do a meta-analysis; this article is about what changes when the inputs are observational.
MOOSE: The Reporting Standard You Are Missing
PRISMA 2020 governs the reporting of any systematic review, but it was built with intervention reviews in mind and underweights the things that matter most in observational synthesis. MOOSE, the Meta-analysis Of Observational Studies in Epidemiology reporting guideline, adds them (Stroup et al., 2000). It is a 35-item checklist across six domains: background, search strategy, methods, results, discussion, and conclusions, and it is listed on the EQUATOR Network.
Two MOOSE items are the ones PRISMA does not stress, and they are exactly where observational reviews fail. Item 20 requires an explicit assessment of confounding, and item 32 requires you to consider alternative explanations for the observed association. Both exist because an observational association has explanations a trial result does not, and a review that ignores them overclaims. There is no revised "MOOSE 2020"; the original remains current, and best practice is to report against MOOSE and PRISMA 2020 together, PRISMA for the overall structure and flow diagram, MOOSE for the observational-specific items.
ROBINS-E: The Risk-of-Bias Tool That Replaces the Star System
Here is the single most common weakness in these reviews. Many still assess study quality with the Newcastle-Ottawa Scale, awarding stars and sometimes summing them into a threshold. The methodological literature has criticized this for over a decade. Stang's evaluation concluded the scale has unknown or even invalid validity for ranking study quality in meta-analyses (Stang, 2010), and a reliability test found agreement between two reviewers applying it was only fair, ranging from poor to substantial across items (Hartling et al., 2013). The field has moved from summary quality scores toward structured, domain-based risk-of-bias assessment that judges each specific result.
For observational studies of exposures, that tool is ROBINS-E, Risk Of Bias In Non-randomized Studies of Exposures, available at riskofbias.info (Higgins et al., 2024). It is distinct from ROBINS-I, which is for non-randomized studies of interventions (Sterne et al., 2016); the two share architecture but answer different questions, so use ROBINS-E when your studies concern an exposure and ROBINS-I when they concern an intervention. We cover the intervention tool and the wider appraisal toolkit in our guides on critically appraising studies and the ROBINS-I v2 update.
ROBINS-E has a defining feature worth understanding before you start: the target trial. Before assessing bias, you specify the hypothetical pragmatic randomized trial that the observational study is trying to emulate, and you then judge bias as the study's deviation from that ideal. Assessment runs across seven bias domains through signaling questions, and each domain and the overall result are rated on a four-level scale: low risk, some concerns, high risk, or very high risk. Because uncontrolled confounding can never be fully excluded in an observational study, the best attainable overall rating for many findings is effectively low risk except for concerns about residual confounding.
Table 1: The Seven ROBINS-E Bias Domains
Domain | What it examines |
|---|---|
1. Confounding | Whether a third factor tied to exposure and outcome distorts the association, and how well it was controlled |
2. Measurement of the exposure | Whether the exposure was classified accurately and independently of the outcome |
3. Selection into the study or analysis | Whether who entered the study or the analysis biased the association |
4. Post-exposure interventions | Whether things that happened after exposure affected the outcome and were mishandled |
5. Missing data | Whether missing exposure, outcome, or confounder data biased the result |
6. Measurement of the outcome | Whether the outcome was ascertained accurately and independently of the exposure |
7. Selection of the reported result | Whether the reported estimate was cherry-picked from multiple analyses or outcomes |
Two practical cautions. ROBINS-E is demanding, and its sibling ROBINS-I is documented as frequently misapplied, with reviewers modifying the scale or understating bias, so use two independent assessors and reconcile. And the risk-of-bias judgments feed directly into GRADE and your summary-of-findings table, so run ROBINS-E before you rate certainty. The screening, extraction, and risk-of-bias workload here is substantial, which is what our screening and data extraction service handles.
Handling the Estimates: Where Observational Pooling Lives or Dies
The statistical choices that distinguish a defensible observational meta-analysis from a rejected one are specific.
Use a random-effects model, almost always. Because observational studies differ systematically, the assumption of one common true effect rarely holds. The standard estimator is DerSimonian-Laird (DerSimonian & Laird, 1986), but it underestimates uncertainty when studies are few, so the recommended refinement is the Hartung-Knapp-Sidik-Jonkman adjustment, which produces more reliable confidence intervals (IntHout et al., 2014).
Pool adjusted estimates, and be explicit about which confounders each study adjusted for. This is the crux. Combining crude estimates from some studies with adjusted estimates from others, or pooling estimates adjusted for wildly different covariate sets, is one of the top reasons these reviews are rejected. Extract the maximally adjusted estimate from each study, tabulate the adjustment set per study, and run sensitivity analyses by adjustment level. Do not combine incompatible effect measures either: odds ratios, risk ratios, and hazard ratios estimate different things and should not be pooled as if interchangeable. Handling this correctly is part of what our meta-analysis service exists to do, and candidates analyzing their own primary data alongside a synthesis can draw on our statistical analysis service.
Report heterogeneity fully. Give tau-squared, I-squared, and a prediction interval (Higgins et al., 2003). The critical point: I-squared is routinely high in observational synthesis, and a high value is not a reason to abandon the analysis or to ignore it. It is a prompt to explore why, through pre-specified subgroups and meta-regression, and to report a prediction interval that shows the real dispersion of effects. What to do when heterogeneity is high is its own subject, which we cover in our guide on heterogeneity in meta-analysis.
For graded exposures, model the trend rather than dichotomizing, using dose-response meta-analysis (Greenland & Longnecker, 1992). A demonstrated dose-response gradient is also a reason to rate certainty up under GRADE.
On publication bias, funnel plots and Egger's test are available (Egger et al., 1997), but interpret them cautiously. Funnel-plot asymmetry in observational data often reflects true heterogeneity rather than publication bias, and the tests are underpowered with fewer than about ten studies. Report them as one input among several, not as a verdict.
Unsure how to pool your adjusted estimates? |
|---|
Reconciling different adjustment sets, choosing the right random-effects model, and reading a high I-squared correctly are where observational meta-analyses go wrong. Send us your extracted estimates, and a statistician will run the pooling, the sensitivity analyses, and the heterogeneity diagnostics properly. Ask about a meta-analysis and get an itemized quote within 2 to 4 business hours, no obligation. |
GRADE for Observational Evidence: The Point Candidates Miss
Under GRADE, observational studies start at low certainty, which many candidates take as a ceiling. It is not. Observational evidence can be rated up, and knowing how is a genuine advantage (Guyatt et al., 2011).
Three conditions allow rating up, provided the evidence is not already rated down for risk of bias, inconsistency, indirectness, imprecision, or publication bias. A large effect: from at least two studies with no plausible confounders, a relative risk above 2 or below 0.5 warrants rating up one level, and above 5 or below 0.2 warrants two levels. A dose-response gradient. And plausible confounding that would work in the opposite direction, where every plausible residual confounder would have reduced the observed effect, so the true effect is likely at least as large as observed.
The practical sequence is to run ROBINS-E first, and if the studies are at low or some-concerns risk and show a large or dose-responsive effect, you can defensibly move certainty from low to moderate or high. This start-low-but-rate-up logic is specific to observational evidence, and it sits inside the wider GRADE framework we explain in our guide on GRADE certainty ratings.
The Workflow, End to End
The defensible pipeline: register the protocol on PROSPERO before extraction, naming your confounders of concern, your risk-of-bias tool, your effect measure, and your subgroup plans. Our protocol and registration service sets that up. Report against PRISMA 2020 and MOOSE together. Assess each result with ROBINS-E across the seven domains, anchored to the target trial. Pool with a random-effects model with the HKSJ adjustment, reporting tau-squared, I-squared, and a prediction interval. Assess small-study effects only with ten or more studies. Feed the ROBINS-E judgments into GRADE, apply the rate-down and rate-up criteria, and build a summary-of-findings table. The write-up that ties this together is where our systematic review writing service and the wider systematic review services come in.
Why These Reviews Get Rejected
The failures are specific and, once named, avoidable.
Table 2: Why Observational Meta-Analyses Get Rejected, and the Fix
Failure | The fix |
|---|---|
No structured risk-of-bias assessment, or NOS used as a score | Assess each result with ROBINS-E across seven domains, two independent assessors |
Pooling crude and adjusted estimates together | Pool maximally adjusted estimates; tabulate each study's adjustment set; sensitivity-analyze by level |
Combining incompatible effect measures (OR, RR, HR) | Pool like with like; convert or separate where appropriate |
Ignoring confounding | Name key confounders a priori (MOOSE item 20); tabulate adjustment per study |
Over-claiming causation from a significant pooled association | State it is an association; discuss alternative explanations (MOOSE item 32) |
Treating high I-squared as ignorable | Explore it with pre-specified subgroups; report a prediction interval |
No MOOSE or PROSPERO registration | Register prospectively; report against MOOSE plus PRISMA 2020 |
Even when the right tool is chosen, it is often misapplied. A methodological review of ROBINS-I use found that a fifth of reviews modified the rating scale, a fifth understated overall risk of bias, and risk of bias was serious or critical in over half of assessments, most often because of confounding. Expect reviewers to scrutinize how you ran the tool, not just whether you named it. Each of these failures is a decision made before writing, in the protocol and the analysis plan, which is why they cannot be edited out at the end.
Frequently Asked Questions
What is the difference between MOOSE and PRISMA?
PRISMA 2020 is the general reporting guideline for any systematic review; MOOSE is specific to meta-analyses of observational studies in epidemiology. They are complementary, not competing. PRISMA governs the overall structure, search reporting, and flow diagram, while MOOSE adds observational-specific items that PRISMA underweights, most importantly, an explicit assessment of confounding and consideration of alternative explanations for the observed association. Report against both.
Which risk-of-bias tool should I use for observational studies?
For observational studies of exposures, use ROBINS-E; for non-randomized studies of interventions, use ROBINS-I. Both are structured, domain-based tools that judge each specific result, and both are preferable to the Newcastle-Ottawa Scale, which the methodological literature has criticized for poor reliability and validity as a summary quality score. ROBINS-E assesses seven bias domains through signaling questions, anchored to a target trial, and rates each result as low risk, some concerns, high risk, or very high risk.
Is the Newcastle-Ottawa Scale still recommended?
It is widely used, and some journals still request it, but it is no longer recommended as your primary or sole tool. Critiques have documented poor inter-rater reliability and questionable validity when its stars are summed into a quality score. Use a domain-based tool like ROBINS-E instead. If a journal requires the Newcastle-Ottawa Scale, use it only as a secondary descriptor, never as a numeric threshold that weights or excludes studies, and pair it with a proper risk-of-bias assessment.
What is the difference between ROBINS-E and ROBINS-I?
They share the same architecture but apply to different study types. ROBINS-E (Exposures) is for observational studies of exposures, such as environmental, occupational, dietary, or behavioral factors that were not assigned. ROBINS-I (Interventions) is for non-randomized studies of interventions, where something was done to participants but not through randomization. Choose the tool based on what your studies examine: an exposure people were subject to, or an intervention that was delivered.
How do I handle adjusted versus crude estimates in a meta-analysis?
Pool adjusted estimates, not crude ones, and never mix the two without justification. Extract the maximally adjusted estimate from each study, then tabulate which confounders each study adjusted for, because studies adjusted for very different covariate sets are not directly comparable. Run sensitivity analyses by adjustment level to show your pooled result is robust. Combining crude and adjusted estimates, or estimates with incompatible adjustment sets, is one of the most common reasons observational meta-analyses are rejected.
Can observational studies be high certainty under GRADE?
Yes. Observational evidence starts at low certainty in GRADE, but it can be rated up. If the evidence is not already downgraded for other reasons, a large effect (a relative risk above 2 or below 0.5 from studies with no plausible confounders, two levels if above 5 or below 0.2), a dose-response gradient, or plausible confounding that would work against the observed effect can each raise certainty. So a well-conducted observational synthesis can reach moderate or even high certainty.
Does a significant pooled result prove causation?
No. A statistically significant pooled estimate from observational studies summarizes an association, not a causal effect. If the contributing studies share the same confounding or bias, meta-analysis averages that bias and narrows the confidence interval around a potentially wrong point, an effect described as spurious precision. Establishing causation requires considering confounding, alternative explanations, dose-response, and the coherence of the wider evidence, which is exactly what MOOSE item 32 and the GRADE framework ask you to do.
You May Also Find Useful
- How to Critically Appraise Studies in a Systematic Review- RoB 2, ROBINS-I V2, GRADE, CASP
- ROBINS-I V2 (2025)
- GRADE Certainty Ratings: Why Your Evidence Was Downgraded and How to Defend Each Domain
- Heterogeneity in Meta-Analysis: What to Do When I² Is High
- How to do a Meta-Analysis: REML, Forest Plots, and GRADE
Pooling Observational Evidence Without Overclaiming
A meta-analysis of observational studies is defensible when it respects what observational data are. Report against MOOSE and PRISMA 2020, so your confounding assessment and alternative explanations are on the record. Assess every result with ROBINS-E, anchored to a target trial, not with a star score. Pool adjusted estimates with compatible adjustment sets, use random effects with the HKSJ adjustment, and report a prediction interval alongside a high I-squared rather than hiding it. Rate certainty with GRADE, and use the rate-up criteria your evidence earns. And never let a narrow confidence interval tempt you into claiming causation that the design cannot support. Do that, and the synthesis answers the exact objections reviewers raise about observational evidence.
If you want a statistician to run that synthesis, or to diagnose why one came back, send us your studies and extracted estimates. You will have an itemized quote within 2 to 4 business hours, no obligation.

