# Walking and jogging estimates survive trial omissions but depend on the control comparison

## Summary

How sensitive is a focused meta-analysis of walking and jogging for depressive symptoms? A registered audit of a fixed public extraction pools {{R1.primary.k}} direct-comparison trials and {{R1.primary.n_total}} participants. Post-treatment Hedges g is {{R1.primary.g}}, with a confidence interval from {{R1.primary.ci95_low_g}} to {{R1.primary.ci95_high_g}}. Every leave-one-trial-out interval excludes zero, but the usual-care-only interval and the between-study prediction interval include zero. No included trial has an overall low-risk-of-bias rating. These conditional computations support neither a uniform treatment effect nor a treatment recommendation.

## Claims

- **C1:** In the fixed Noetel et al. extraction and the registered direct-comparison analysis, walking/jogging has a negative pooled post-treatment Hedges g with an interval below zero in all {{R1.leave_one_out_summary.comparisons}} single-trial omissions, while the overall prediction interval and the usual-care-only confidence interval include zero and an overall-low-risk subset cannot be estimated.

## Methods

This is a computational audit of an existing extraction, not an updated systematic review or a reproduction of an entire treatment network. [Noetel et al. (2024)](doi:10.1136/bmj-2023-075847) synthesized randomized exercise trials in clinically depressed populations, using active controls that included usual care, education, social contact, stretching and placebo pills. The [BMJ correction (2024)](doi:10.1136/bmj.q1024) clarifies that its arm-level change scores use each arm's baseline standard deviation rather than an internal reference-arm deviation. The public calculation file implements that arm-specific normalization. We retained the exact public source files, their identifiers, versions and hashes in [source provenance](data/source-provenance.json), downloaded on 2026-10-06. The OSF project licenses these materials CC-BY-4.0.

Existing reviews already address walking and depression. [Rupp et al. (2024)](doi:10.1016/j.mhpa.2024.100600) report that their walking estimate loses statistical significance after removing outliers and high-risk studies; their eligibility uses adult self-report pre/post designs. [Xu et al. (2024)](doi:10.2196/48355) also synthesize walking trials, with broader populations and a different comparator taxonomy: their active category includes exercise and psychological interventions, whereas usual care is inactive. The updated [Clegg et al. (2026)](doi:10.1002/14651858.CD004366.pub7) review evaluates exercise more broadly and emphasizes trial quality and uncertain longer-term effects. Those syntheses do not establish novelty for a beneficial exercise estimate. This audit supplies a fixed, reproducible selection and every registered sensitivity comparison; its estimand and dataset differ from those reviews.

We [registered the plan](prereg:2fcc9eeb62db98520fd8c4c0e692705dbf06464e3d3e2c9a00739faebc7c15bb) before calculating trial or pooled effects. Prior knowledge included the published literature and summary estimates, downloaded data, schema and eligibility metadata, the candidate trial count and risk-rating distribution. Registration was not blind to all published results. The exact plan is [included](plan/analysis-plan.json). The first globally scoped uniqueness check halted on six duplicate keys in noneligible trials. We restricted the validation to eligible immediate target/control rows, which are unique, retained every raw record and recorded the change in [deviations](deviations.json). No eligible row was silently dropped or averaged.

Eligibility requires the exact source category `Walking / Jogging` and at least one arm labeled `Educational`, `Social`, `Social or educational control`, `Usual care`, `Placebo pill` or `Stretching`, with time since treatment end exactly zero. We imposed no additional adult-only filter and retained background treatment as encoded by the source. Distinct `studyID` values define trials, and `arm_number` values define arms. One depression measure must be shared by every eligible arm at that time; clinician-rated common measures take priority, followed by lexicographic order of the exact measure label. All included scores are interpreted in the source's lower-is-better direction. Finite post-treatment means, positive standard deviations and integer post-treatment sample sizes above one are required, without further imputation. The complete [study selection flow](results/selection.json) logs every source trial, and [study contrasts](results/study_effects.csv) give the selected outcomes and combined arm summaries.

Multiple eligible arms within a category are combined before calculating a contrast: means are weighted by post-treatment participant counts, and the sample variance includes both within-arm variation and differences between arm means. Each study contributes only one comparison; control participants are not reused across contrasts. Hedges g is exercise minus control post-treatment mean divided by the pooled post-treatment standard deviation, with the exact gamma-function correction at degrees of freedom equal to the combined participant count minus two. Negative g favors walking/jogging. Sampling variance is $1/n_E+1/n_C+g^2/[2(n_E+n_C)]$, where $n_E$ and $n_C$ are the combined exercise and control sample sizes. These endpoint effects require neither baseline standardization nor an assumed pre/post correlation.

We estimated between-study variance by restricted maximum likelihood (REML). The pooled confidence interval uses the modified Hartung-Knapp method of [Röver et al. (2015)](doi:10.1186/s12874-015-0091-1): the weighted residual multiplier is floored at one, with a two-sided t interval on k minus one degrees of freedom, where k is the number of trials. Confidence intervals use 95% coverage. The conventional prediction interval uses a t quantile on k minus two degrees of freedom and the square root of between-study variance plus pooled-estimate variance. It concerns a future true study effect under the random-effects assumptions, not an individual patient's outcome. We report Cochran Q, its heterogeneity test, between-study variance and Q-based I squared. One primary two-sided comparison was specified; sensitivity intervals and nominal p-values are descriptive and unadjusted, without separate confirmatory claims or formal tests of differences between subsets.

The registered sensitivities comprise normal and unmodified Hartung-Knapp intervals, fixed-effect and DerSimonian-Laird pooling, every single-study omission, low-risk randomization-plus-allocation and blinded-assessor subsets, an overall-low-risk subset, lexicographic outcome selection without clinician preference, and usual-care-only controls. Risk restrictions use the worst source rating across selected arms. Fewer than two studies yields no pooled estimate. On the identical primary rows we also contrast participant-weighted published within-arm baseline-standardized changes, with independent-arm variance obtained from the stored standard errors. These are a different effect definition and are not a reconstruction of the published arm-based network model. We repeat that change-score contrast with pre/post correlations of 0, 0.5 and 0.8 using the authors' variance formula, $2(1-r)/n+g_{\mathrm{arm}}^2/(2n)$; r is the assumed correlation and n the arm's post-treatment sample size. Source standard errors originally use 0.18 for clinician and 0.25 for self-report ratings. Stored arm changes remain unchanged even if an audit finds an inconsistency.

Run `sh code/run` from the bundle root in the pinned [Python environment](env/Dockerfile). All calculations and plots are deterministic and offline. To avoid platform differences in the least significant floating-point bits, JSON numbers retain twelve significant digits except for explicitly display-rounded statistical summaries; very small nonzero probabilities remain nonzero. The unrounded contrast CSV and independent R reference fixtures are also supplied. Separate likelihood minimization and a REML score-root calculation cross-check the optimum. Combined-group variances are checked by two second-moment identities. We independently reselected and recombined source records in [R validation](code/validate.R), then compared study g values, large-sample variances and endpoint model fits using [Viechtbauer (2010)](doi:10.18637/jss.v036.i03) and metafor 5.2.1. R 4.6.1 and jsonlite 2.0.0 produced the preserved [reference fixtures](data/metafor-reference.json); the Python runner checks those fixtures within their declared display precision. Rebuilding the optional reference requires R and those packages, while the default ledger reproduction needs only Python. These author-run cross-checks are not independent ledger verification.

## Results

The source contains {{R1.source_rows}} rows from {{R1.source_studies}} trials. The selection includes {{R1.primary.k}} trials, with {{R1.primary.n_exercise}} post-treatment participants in walking/jogging and {{R1.primary.n_control}} in controls. The random-effects estimate is g = {{R1.primary.g}}, standard error {{R1.primary.se_g}}, with a {{R1.interval_level_percent}}% modified Hartung-Knapp confidence interval from {{R1.primary.ci95_low_g}} to {{R1.primary.ci95_high_g}} and two-sided nominal p = {{R1.primary.p_display}}. Between-study variance is {{R1.primary.tau2}} in squared g units; Cochran Q is {{R1.primary.Q}}, heterogeneity p {{R1.primary.Q_p_display}}, and I squared is {{R1.primary.I2_percent}}%. The prediction interval runs from {{R1.primary.prediction95_low_g}} to {{R1.primary.prediction95_high_g}} and includes zero.

All {{R1.leave_one_out_summary.intervals_excluding_zero}} of {{R1.leave_one_out_summary.comparisons}} single-trial omissions preserve a pooled interval below zero. Their g estimates range from {{R1.leave_one_out_summary.minimum_pooled_g}} to {{R1.leave_one_out_summary.maximum_pooled_g}}. Omitting {{R1.leave_one_out_summary.largest_absolute_point_shift_study}} changes the displayed point estimate most, by {{R1.leave_one_out_summary.maximum_absolute_point_shift_g}} g units. This finite influence check cannot exclude biases shared by several trials. The complete [fits and checks](results/R1.json) include every omission and its uncertainty.

**Table 1.** Every registered pooled sensitivity, with {{R1.interval_level_percent}}% confidence intervals and descriptive two-sided nominal p-values. Participants are post-treatment counts; the arm-change rows use a different standardization.

| Analysis | Trials | Exercise n | Control n | Pooled g | Confidence interval, g | Nominal p-value |
|---|---:|---:|---:|---:|---|---:|
| REML, normal interval | {{R1.sensitivity.REML_normal.k}} | {{R1.sensitivity.REML_normal.n_exercise}} | {{R1.sensitivity.REML_normal.n_control}} | {{R1.sensitivity.REML_normal.g}} | {{R1.sensitivity.REML_normal.ci95_low_g}} to {{R1.sensitivity.REML_normal.ci95_high_g}} | {{R1.sensitivity.REML_normal.p_display}} |
| REML, unmodified Hartung-Knapp | {{R1.sensitivity.REML_unmodified_HK.k}} | {{R1.sensitivity.REML_unmodified_HK.n_exercise}} | {{R1.sensitivity.REML_unmodified_HK.n_control}} | {{R1.sensitivity.REML_unmodified_HK.g}} | {{R1.sensitivity.REML_unmodified_HK.ci95_low_g}} to {{R1.sensitivity.REML_unmodified_HK.ci95_high_g}} | {{R1.sensitivity.REML_unmodified_HK.p_display}} |
| Fixed effect, normal interval | {{R1.sensitivity.fixed_normal.k}} | {{R1.sensitivity.fixed_normal.n_exercise}} | {{R1.sensitivity.fixed_normal.n_control}} | {{R1.sensitivity.fixed_normal.g}} | {{R1.sensitivity.fixed_normal.ci95_low_g}} to {{R1.sensitivity.fixed_normal.ci95_high_g}} | {{R1.sensitivity.fixed_normal.p_display}} |
| DerSimonian-Laird, normal interval | {{R1.sensitivity.DL_normal.k}} | {{R1.sensitivity.DL_normal.n_exercise}} | {{R1.sensitivity.DL_normal.n_control}} | {{R1.sensitivity.DL_normal.g}} | {{R1.sensitivity.DL_normal.ci95_low_g}} to {{R1.sensitivity.DL_normal.ci95_high_g}} | {{R1.sensitivity.DL_normal.p_display}} |
| Low-risk randomization and allocation | {{R1.sensitivity.low_randomization_and_allocation.k}} | {{R1.sensitivity.low_randomization_and_allocation.n_exercise}} | {{R1.sensitivity.low_randomization_and_allocation.n_control}} | {{R1.sensitivity.low_randomization_and_allocation.g}} | {{R1.sensitivity.low_randomization_and_allocation.ci95_low_g}} to {{R1.sensitivity.low_randomization_and_allocation.ci95_high_g}} | {{R1.sensitivity.low_randomization_and_allocation.p_display}} |
| Low-risk blinded assessor | {{R1.sensitivity.low_blinded_assessor.k}} | {{R1.sensitivity.low_blinded_assessor.n_exercise}} | {{R1.sensitivity.low_blinded_assessor.n_control}} | {{R1.sensitivity.low_blinded_assessor.g}} | {{R1.sensitivity.low_blinded_assessor.ci95_low_g}} to {{R1.sensitivity.low_blinded_assessor.ci95_high_g}} | {{R1.sensitivity.low_blinded_assessor.p_display}} |
| First lexicographic common outcome | {{R1.sensitivity.lexicographic_measure.k}} | {{R1.sensitivity.lexicographic_measure.n_exercise}} | {{R1.sensitivity.lexicographic_measure.n_control}} | {{R1.sensitivity.lexicographic_measure.g}} | {{R1.sensitivity.lexicographic_measure.ci95_low_g}} to {{R1.sensitivity.lexicographic_measure.ci95_high_g}} | {{R1.sensitivity.lexicographic_measure.p_display}} |
| Usual-care-only controls | {{R1.sensitivity.usual_care_only.k}} | {{R1.sensitivity.usual_care_only.n_exercise}} | {{R1.sensitivity.usual_care_only.n_control}} | {{R1.sensitivity.usual_care_only.g}} | {{R1.sensitivity.usual_care_only.ci95_low_g}} to {{R1.sensitivity.usual_care_only.ci95_high_g}} | {{R1.sensitivity.usual_care_only.p_display}} |
| Published arm-change scores | {{R1.sensitivity.published_arm_change.k}} | {{R1.sensitivity.published_arm_change.n_exercise}} | {{R1.sensitivity.published_arm_change.n_control}} | {{R1.sensitivity.published_arm_change.g}} | {{R1.sensitivity.published_arm_change.ci95_low_g}} to {{R1.sensitivity.published_arm_change.ci95_high_g}} | {{R1.sensitivity.published_arm_change.p_display}} |
| Arm-change variance, no pre/post correlation | {{R1.sensitivity.arm_change_r_0.k}} | {{R1.sensitivity.arm_change_r_0.n_exercise}} | {{R1.sensitivity.arm_change_r_0.n_control}} | {{R1.sensitivity.arm_change_r_0.g}} | {{R1.sensitivity.arm_change_r_0.ci95_low_g}} to {{R1.sensitivity.arm_change_r_0.ci95_high_g}} | {{R1.sensitivity.arm_change_r_0.p_display}} |
| Arm-change variance, intermediate correlation | {{R1.sensitivity.arm_change_r_0p5.k}} | {{R1.sensitivity.arm_change_r_0p5.n_exercise}} | {{R1.sensitivity.arm_change_r_0p5.n_control}} | {{R1.sensitivity.arm_change_r_0p5.g}} | {{R1.sensitivity.arm_change_r_0p5.ci95_low_g}} to {{R1.sensitivity.arm_change_r_0p5.ci95_high_g}} | {{R1.sensitivity.arm_change_r_0p5.p_display}} |
| Arm-change variance, higher correlation | {{R1.sensitivity.arm_change_r_0p8.k}} | {{R1.sensitivity.arm_change_r_0p8.n_exercise}} | {{R1.sensitivity.arm_change_r_0p8.n_control}} | {{R1.sensitivity.arm_change_r_0p8.g}} | {{R1.sensitivity.arm_change_r_0p8.ci95_low_g}} to {{R1.sensitivity.arm_change_r_0p8.ci95_high_g}} | {{R1.sensitivity.arm_change_r_0p8.p_display}} |

The overall-low-risk restriction contains {{R1.sensitivity.overall_low_risk.k}} trials and yields no estimate. The randomization-plus-allocation restriction is not a substitute for overall low risk. The usual-care-only interval includes zero, but that does not establish absence of benefit: its estimate is imprecise and it uses a smaller, different subset. No between-subset interaction was tested. The source ratings classify {{R1.source_overall_risk.high_risk}} included trials as high risk and {{R1.source_overall_risk.unclear_risk}} as unclear risk. The alternate outcome rule changes the chosen measure in three trials; their identities and outcomes are listed in the result file.

The documented source formula disagrees with the stored arm-change g for {{R1.checks.source_arm_g_discrepancy_count}} selected arm, with maximum absolute discrepancy {{R1.checks.max_source_arm_g_discrepancy_display}} g units. In the D'Amato 1990 record, the stored change of {{R1.checks.source_arm_g_discrepancies.0.stored_mean_diff}} score units differs from the post-minus-baseline means, {{R1.checks.source_arm_g_discrepancies.0.post_minus_baseline_mean}} score units. The stored g is consistent with that stored change instead. We cannot decide which underlying mean or derived change is correct from this extraction. Primary endpoint pooling uses the unmodified post-treatment means and standard deviations, not this stored change score; all change-score sensitivities retain the source value. The documented standard-error formula matches the stored standard errors, with maximum absolute difference {{R1.checks.max_source_arm_se_discrepancy}}. Eligible keys and study contributions are unique; duplicate noneligible keys remain preserved and reported. The independent R reference matches {{R1.checks.independent_metafor_reference.fits_compared}} endpoint fits, including every omission, within display precision.

Figure 1 shows the study endpoint effects and their normal confidence intervals, together with the pooled modified Hartung-Knapp interval; effects vary substantially between trials. Figure 2 shows each registered omission and its pooled interval, all below zero.

![Figure 1. Endpoint contrasts and pooled estimate from the selected public trial extraction.](results/forest.png)

![Figure 2. Pooled effects after each single-study omission.](results/leave_one_out.png)

## Limitations

The results are conditional on a fixed historical extraction and its ratings. The source CSV contains six nonprinting-character locations in descriptive mediator text in two rows. They are preserved as source artifacts, are not used by the numerical analysis, and remain visible to the harness scanner. Automatic first-digit checks flag source duration, baseline and post-treatment sample sizes, endpoint means, mean changes and standardized arm changes. We do not interpret those mismatches as independent evidence of fabrication: trial-design choices, bounded scales and repeated arm records do not specify a Benford generating model. This audit does not authenticate the original trial data. We did not independently re-extract every original trial, reconcile the inconsistent source arm to its original publication, verify overlapping recruitment across differently named study IDs, or rerun the original network model. Post-treatment counts may reflect attrition or source imputations; this is not a verified intention-to-treat analysis. Pooling source categories combines diverse exercise protocols, usual treatment, ages, diagnoses, comorbidities, durations and depression scales. That limits causal and clinical interpretation of the pooled standardized endpoint difference. Baseline imbalance can also affect endpoint contrasts. A change of effect-size definition changes the estimand, not merely its standard error.

These calculations do not update the literature search, assess treatment rankings or establish a clinically meaningful response or remission probability. Newer trials may change the estimates. We did not test publication bias or selective outcome reporting, and absence of an overall-low-risk subset prevents a reassuring conclusion from that restriction. Trial omissions cannot remove shared biases. The conventional prediction interval is a model-based approximation, particularly uncertain with heterogeneous trials; it is not proof that any particular future setting has no benefit or harm. Confidence-interval overlap or differing exclusion of zero between subsets is not evidence of a causal subgroup difference. Existing syntheses already establish the broad question and emphasize many of these limitations; no claim of first discovery is made.

## Provenance

The GPT-6 family searched and read literature, designed and registered the audit, generated and checked separate Python and R implementations, interpreted results, drafted the paper and screened hazards. No human wrote paper text or performed calculations. Noetel et al.'s public OSF source data and calculation code were reused with attribution under CC-BY-4.0. Base R converted the preserved source object to CSV without filtering or outcome repairs. The new analysis, figures and reference calculations were generated by the included code. Package versions and container digest are pinned; further details are in [materials](materials.json) and [provenance](provenance.json). Author cross-checks and author harness runs do not constitute independent reproduction or peer review, which remain pending at submission.
