# Methods review

Scope: paper, twelve claim statements, registered plan and deviations, extraction metadata, shared statistical/model/design/IO code, selected association implementations and their stored results, results tables, and environment were examined. Two metadata downloads were fetched and their SHA-256 and byte counts verified. I did not execute or import any supplied program, reproduce the participant-level fits, inspect every participant record, or independently retrieve all forty source articles. This is a methods assessment with an independent arithmetic audit of supplied fit summaries, not a reproduction or an independent confirmation of all inferred original-paper coding decisions. No publisher was identified or searched for; model-family provenance alone does not identify an operator.

The independent aggregate audit agrees with all three BH-adjustment families to about 2e-15, and with the principal counts (14 replicated, 13 informative, 10 informative replications, 16 inside published intervals, 16 same-sign unadjusted positives, 3 published-effect differences, 0 harmonized heterogeneity positives, 39 inside original intervals, 36 papers with labelled headline departures). Only code/run differs from its registered counterpart; adding code/report.R is consistent with the declared reporting deviation. The source inventory has 412 rows. A fitting reproduction in the declared container remains a separate task.

## Required changes

1. C7 needs a less categorical interpretation or a sensitivity count. The program counts investigator-entered departure labels; matching a rounded published coefficient or interval with one reconstruction does not uniquely establish which computation its original authors performed. Some files themselves describe unresolved samples, conflicting prose and table footnotes, multiple plausible coding choices, or substitutions for imputation. For example row040 cannot identify the original imputed sample, row303's chosen full model does not reach the published headline point estimate, and row319 retains an unexplained discrepancy. Separate direct internal contradictions from inferred implementation choices and unresolved mismatches, report the evidence strength, and provide a confirmed-only count. The present aggregate proves the number of labelled findings, not that each label establishes the original implementation. This is a major issue for C7's assertion about what the original analyses computed, rather than evidence that the cross-cycle aggregate calculation is wrong.

2. C1's standalone wording should identify the eligible, accessible, constructible subset of the 341-paper frame. Eligibility and a first-40 stopping rule make this a random sample within those restrictions, not an unconditional sample of all 341 papers. The paper documents its exclusions adequately. Keep population conclusions within that scope.

3. C2's power is a plug-in, unadjusted alpha=0.05 calculation at the published effect, using the new-cycle estimated standard error. It is not the power of the BH-defined replication decision, whose cutoff depends on the other tests. The plan explicitly defines the calculation, so no undeclared change is alleged. State this distinction beside the informative-subset result; that post-data subset is not a separately randomized group, and selected published effects can overstate the effect used for planning.

4. The Clopper-Pearson and order-statistic intervals are mathematically conventional under their sampling models, but associations reuse participants, outcomes and exposure constructs. Their fitted effects and replication decisions need not be independent. Sampling papers without replacement does not remove shared-survey estimation dependence. C1-C3, C6 and C8 should qualify interval coverage or include a dependence-aware sensitivity analysis. The point summaries remain useful descriptive quantities.

5. C4/C8 use normal approximations for differences, while individual weighted fits have 15 design degrees of freedom. Treat the difference tests as approximate and show a finite-df sensitivity if they support strong claims. Non-significant heterogeneity is not evidence of equivalence. C11/C12 describe opposite point-estimate signs versus the original estimates; neither is a significant opposite-sign association after the primary BH correction. The paper already distinguishes these concepts, and that distinction should persist in standalone statements.

6. C6's reproduction criterion is inclusion of the reconstructed estimate within the original published 95% interval, not equality of estimate, sample, model or interval. Say 'meets the declared interval-inclusion criterion' when interpreting 39/40. The extensive alternative-model audit is valuable, but interval inclusion alone cannot validate an implementation or settle a reporting contradiction.

7. C10 should explicitly retain the sign of the published coefficient: results/R2.json gives -1.00, and the paper's bound placeholders preserve that sign. 'Against a published 1.00 g/L' can be read as +1.00. Write the signed coefficient or state that the original also found lower albumin. Its comparison is to a negative original coefficient, not a positive one. The submitted table and computed difference are otherwise consistent.

## Claim assessments

C1 minor_issues: operational replication count agrees; clarify restricted target population and interval assumptions.
C2 minor_issues: the declared calculation is implemented; clarify unadjusted power versus BH decision and plug-in selection.
C3 minor_issues: median agrees; qualify dependence and that it mixes heterogeneous estimands on their analysis scales.
C4 minor_issues: count agrees; qualify normal difference-test approximation and distinction from primary reversal significance.
C5 sound: descriptive interval-inclusion and sign counts agree under the declared rules.
C6 minor_issues: interval-inclusion count agrees; qualify the meaning of reproduced and the inference about original coding.
C7 major_issues: label counting is correct but the categorical interpretation is not established by this inferential reconstruction alone.
C8 minor_issues: count agrees; non-rejection is not equivalence and interval/difference-test assumptions require qualification.
C9 sound: both the count and identities of replicated associations agree across the stored paper-weight sensitivity; the reproduced subset check agrees.
C10 minor_issues: explicitly fix the published coefficient's sign in the standalone statement; computation agrees with the negative value.
C11 sound: stored estimate and published-difference test support the numerical, noncausal cross-cycle comparison; no significant primary reversal is asserted.
C12 sound: same qualification; small sample uncertainty is disclosed and an opposite point-estimate sign is distinguished from primary significant reversal.

Significance: moderate for C1-C9 as a reusable, prespecified meta-research audit of an existing literature; minor for C10-C12 as specific cross-sectional descriptive comparisons. These are judgments of contribution, not independence or causal evidence.

## External primary-source context

NCHS, Brief Overview of Sample Design, Nonresponse Bias Assessment, and Analytic Guidelines for NHANES August 2021-August 2023, first published 2024-09-20:
https://wwwn.cdc.gov/nchs/nhanes/continuousnhanes/overviewbrief.aspx?Cycle=2021-2023

The official guidance confirms new sampling and dietary-interview modes, lower response rates, phlebotomy weights, and caution in cross-cycle interpretation. It prefers the smallest appropriate component's weights and discusses dietary/blood combinations. The supplied design code's component-specific weights, survey strata and PSUs, and domain analysis are reasonable choices for weighted fits. Inferred unweighted replications reproduce the reported analysis choice; they should not be presented as nationally representative causal estimates. Complete-case substitution and unavailable covariates can change the estimand; the harmonized-original decomposition is a useful disclosed sensitivity, not a full correction for selection or measurement changes.

## Integrity and reproducibility

The assignment's deterministic integrity arrays are empty. The paper has the required sections and bound result placeholders, and no numerical-integrity flag requires dismissal or escalation. Data sources and hashes are documented, compressed public-use files are bundled or offered by authenticated downloads, and the environment pins R and package versions. The participant-level dataset and forty source texts were not independently reanalyzed here, so no claim of complete numerical reproduction, original-article verification, or universal data provenance verification is made. Public-use NHANES data are treated under their released terms; no participant identification was attempted.
