# Single-factor NHANES findings: a third replicated in new data, most tests lacked power, and most analyses differ from their description

## Summary

Many papers relate one NHANES variable to one health condition without correcting for multiple testing ([Suchak et al. (2025)](doi:10.1371/journal.pbio.3003152)). We sampled {{R1.replication.associations}} of their associations, reproduced each on its paper's cycles, and, under a [registered plan](prereg:66c14e8df19aafb51c6890eb73fc46addc6eb292998aa04c081f412f2df6e299), re-ran it on August 2021–August 2023. There, {{R1.replication.replicated.k}} replicated; most tests lacked power, but of the {{R1.replication.informative}} that had it, {{R1.replication.replicated_informative.k}} replicated. Effects were a median {{R1.replication.ratio.median}} times their published size, and {{R1.replication.differs_from_published.k}} differed significantly from it. In {{R1.reproduction.departures.affecting_headline.k}} papers the computation that reproduces the published estimate differs from the paper's description, as the paper itself shows in {{R1.reproduction.departures.shown_by_paper.k}} and recomputation in {{R1.reproduction.departures.identified_by_data.k}} more. Their methods sections poorly describe what their estimates measure, and one cycle can test few of the associations.

## Claims

- **C1:** Of the {{R1.replication.replicated.n}} associations sampled from those eligible, {{R1.replication.replicated.k}} replicated in NHANES August 2021–August 2023, with the published sign and a significant corrected p-value: a share of {{R1.replication.replicated.share}}, with an exact confidence interval from {{R1.replication.replicated.ci.0}} to {{R1.replication.replicated.ci.1}} that treats the associations as independent.
- **C2:** At the nominal significance level, which bounds the power of the corrected replication decision from above, only {{R1.replication.informative}} of the replication tests had the planned power to detect the published effect, and {{R1.replication.replicated_informative.k}} of them replicated (a share of {{R1.replication.replicated_informative.share}}, from {{R1.replication.replicated_informative.ci.0}} to {{R1.replication.replicated_informative.ci.1}}); at the Bonferroni level, which bounds it from below, {{R1.replication.informative_bonferroni}} had it, and {{R1.replication.replicated_informative_bonferroni.k}} of them replicated.
- **C3:** On the analysis scale, the 2021–2023 effects are a median {{R1.replication.ratio.median}} times the published ones ({{R1.replication.ratio.ci.0}} to {{R1.replication.ratio.ci.1}}) over all associations, and {{R1.replication.ratio_informative.median}} times ({{R1.replication.ratio_informative.ci.0}} to {{R1.replication.ratio_informative.ci.1}}) over the informative ones, with distribution-free intervals that treat the associations as independent.
- **C4:** {{R1.replication.differs_from_published.k}} of the 2021–2023 estimates differ significantly from the published ones after correction: {{R1.replication.differs_by_direction.smaller.k}} smaller than published, {{R1.replication.differs_by_direction.larger.k}} larger, and {{R1.replication.differs_by_direction.opposite_sign.k}} on the other side of the null; with t tests on the fits' design degrees of freedom, {{R1.replication.differs_from_published_t.k}} differ.
- **C5:** In 2021–2023, {{R1.replication.in_published_ci.k}} estimates fall inside the published confidence interval, {{R1.replication.same_sign_p05.k}} have the published sign with an unadjusted p-value below the planned threshold, and {{R1.replication.reversed.k}} have the opposite sign with a significant corrected p-value.
- **C6:** On the cycles each paper analyzed, our code meets the registered reproduction criterion, its estimate inside the published interval, for {{R1.reproduction.reproduced.k}} of the published estimates, with a median ratio of our estimate to the published one of {{R1.reproduction.ratio_original.median}} ({{R1.reproduction.ratio_original.ci.0}} to {{R1.reproduction.ratio_original.ci.1}}). The one it doesn't reproduce modeled the absence of asthma while reporting the odds of asthma: coded as stated it gives {{R2.row257.original.estimate}}, and reversed, {{R2.row257.variants.0.estimate}}.
- **C7:** In {{R1.reproduction.departures.affecting_headline.k}} papers, the investigators recorded at least one departure bearing on the headline estimate, where the computation that reproduces it differs from the paper's description. In {{R1.reproduction.departures.shown_by_paper.k}} of them the paper itself shows one, two of its parts disagreeing or its own numbers ruling out what it describes; in {{R1.reproduction.departures.identified_by_data.k}} more, recomputation from the public files shows one, a computation other than the one described being the one that reproduced the published numbers.
- **C8:** {{R1.replication.heterogeneous.k}} associations differ significantly, after correction, between 2021–2023 and our harmonized analysis of the paper's own cycles, with z tests, and {{R1.replication.heterogeneous_t.k}} with t tests on the design degrees of freedom. This does not show the estimates equal: the median ratio of the two is {{R1.replication.ratio_own.median}} ({{R1.replication.ratio_own.ci.0}} to {{R1.replication.ratio_own.ci.1}}).
- **C9:** The replication count is the same among the reproduced associations ({{R1.replication.replicated_reproduced.k}} of {{R1.replication.replicated_reproduced.n}} replicated) and with each paper's own survey weight in place of the 2021–2023 subsample weights ({{R1.replication.replicated_paper_weight.k}} replicated).
- **C10:** In 2021–2023, the coefficient of serum albumin on workday sleep of five hours or less, against more than seven and up to eight, is {{R2.row284.replication.estimate}} g/L ({{R2.row284.replication.low}} to {{R2.row284.replication.high}}), against the published {{R2.row284.published.estimate}} g/L ({{R2.row284.published.low}} to {{R2.row284.published.high}}): a smaller decrement than published, with an interval that includes zero, and a significant difference from it (corrected p {{R2.row284.difference_q}}, or {{R2.row284.difference_q_t}} by t test). On the paper's own cycles, whose albumin analyzers differ, a cycle term in its model gives {{R2.row284.variants.9.estimate}} g/L.
- **C11:** In 2021–2023, the odds ratio of diabetes on the systemic immune-inflammation index, on the paper's scale, is {{R2.row303.replication.estimate}} ({{R2.row303.replication.low}} to {{R2.row303.replication.high}}), against the published {{R2.row303.published.estimate}} ({{R2.row303.published.low}} to {{R2.row303.published.high}}): a significant difference (corrected p {{R2.row303.difference_q}}), the later estimate below one.
- **C12:** In 2021–2023, the odds ratio of depression per unit of the triglyceride-glucose index is {{R2.row311.replication.estimate}} ({{R2.row311.replication.low}} to {{R2.row311.replication.high}}), against the published {{R2.row311.published.estimate}} ({{R2.row311.published.low}} to {{R2.row311.published.high}}): a significant difference by the registered z test (corrected p {{R2.row311.difference_q}}) but not by a t test on the design degrees of freedom (corrected p {{R2.row311.difference_q_t}}), the later estimate below one but its interval wide.
- **C13:** Two independent codings of how each of the {{R1.reproduction.departures.headline_departures}} departures bearing on a headline estimate is established agreed on {{R1.reproduction.departures.coding.agree.k}} (Cohen's kappa {{R1.reproduction.departures.coding.kappa}}). As settled, {{R1.reproduction.departures.by_evidence.paper.k}} are shown by the paper itself, {{R1.reproduction.departures.by_evidence.data.k}} by recomputation that identifies another computation, and {{R1.reproduction.departures.by_evidence.unresolved.k}} are unresolved, the computation described ruled out but the one used not identified.

## Methods

**Design.** A replication of published estimates in a later, independent national sample, analyzed as the papers analyzed theirs. The plan, its eligibility rules, the extraction of every sampled paper, the analysis code, and the coding check on the papers' own cycles were fixed and [registered](prereg:66c14e8df19aafb51c6890eb73fc46addc6eb292998aa04c081f412f2df6e299) on 2026-10-06, before any file of NHANES August 2021–August 2023 was downloaded. The registered files are in [plan/](plan/plan.md), unchanged; the plan says what was settled when, and `deviations.json` lists what was done after registration. This version corrects the first in answer to its reviews: the classification of each departure's evidence, the tests on finite degrees of freedom, and power at the Bonferroni level below were added after the results were known. To judge which variables 2021–2023 has, we read only its documentation, whose codebook pages show each variable's marginal counts but no association between variables.

**Sample.** The frame is Table A of S1 Data of [Suchak et al. (2025)](doi:10.1371/journal.pbio.3003152): 341 papers, published from 2014 to 2024, each relating one NHANES predictor to one health condition (`data/frame.csv`). Its rows were shuffled once in R 4.6.1 with `set.seed(20261006, kind = "Mersenne-Twister", normal.kind = "Inversion", sample.kind = "Rejection")` (`code/sample.R`, which writes `results/order.csv`), and papers were judged in that order until 40 were eligible. A paper was eligible if it was not retracted (E0), its full text was in Europe PMC (E1), it analyzed continuous NHANES (E2), its headline association was cross-sectional (E3), the headline was a significant whole-population estimate from a generalized linear model with a 95% confidence interval (E4), and its exposure, outcome, and defining population could be built from 2021–2023's public files with the same measurement or question (E5). The headline is the first significant estimate the abstract reports for the paper's predictor and condition, its most-adjusted model, and for ordered categories its highest against its lowest; [the plan](plan/plan.md) gives the rules in full. Judging ran through rank 206: 2 papers were retracted, 53 had no full text in Europe PMC, 3 did not analyze continuous NHANES, 4 had no cross-sectional headline, 8 had no qualifying headline, and 96 could not be built from 2021–2023 (85 judged from Table A's labels and abstract, 11 from the full text). Each decision is in [plan/eligibility.csv](plan/eligibility.csv).

The 40 sampled associations, by Table A's labels, are: serum vitamin D concentrations and osteoarthritis ([Yu et al. (2023)](doi:10.3389/fnut.2023.1016809)); systemic immune-inflammation index and stroke ([Liu et al. (2024a)](doi:10.31083/J.RCM2504130)); dietary inflammatory index and stroke ([Mao et al. (2024)](doi:10.1186/s12889-023-17556-w)); the non-HDL to HDL cholesterol ratio and gallstones ([Cheng et al. (2024a)](doi:10.1186/s12944-024-02262-2)); blood pressure and depression ([Zhang et al. (2024a)](doi:10.3389/fpsyt.2024.1433990)); non-HDL cholesterol and depression ([Zhu et al. (2023)](doi:10.3389/fpsyt.2023.1274648)); dietary inflammatory index and hyperuricemia ([Wang et al. (2023)](doi:10.3389/fnut.2023.1218166)); caffeine intake and obesity ([Liu and Cui (2024)](doi:10.1371/journal.pone.0300566)); sleep health and blood pressure ([Su et al. (2022)](doi:10.1038/s41598-022-05124-y)); red blood cell distribution width and coronary heart disease ([Zhang et al. (2024b)](doi:10.1097/MD.0000000000037315)); heavy metals exposure and metabolic-associated fatty liver conditions ([Tang et al. (2024)](doi:10.3389/fpubh.2024.1280163)); serum ferritin levels and the metabolic score for insulin resistance ([Hao et al. (2022)](doi:10.3389/fmed.2022.925344)); weight-adjusted-waist index and metabolic-associated fatty liver conditions ([Hu et al. (2023)](doi:10.1186/s40001-023-01205-4)); blood manganese and liver stiffness ([Han et al. (2023)](doi:10.1186/s40001-022-00977-5)); the neutrophil to HDL cholesterol ratio and metabolic-associated fatty liver conditions ([Lu et al. (2024)](doi:10.1186/s12876-024-03394-6)); blood cadmium levels and depression ([Ji and Wang (2024)](doi:10.1265/ehpm.24-00050)); usual source of care and blood pressure ([Dinkler et al. (2016)](doi:10.1093/ajh/hpw010)); visceral adiposity index and chronic kidney disease ([Peng et al. (2023)](doi:10.1016/j.pmedr.2023.102306)); dietary fiber intake and stroke ([Dong and Yang (2022)](doi:10.3389/fnut.2022.936926)); sleep health and visceral adiposity index ([Liu et al. (2024b)](doi:10.1136/bmjopen-2023-082601)); ovariectomy-reduced hormones and depression ([Chen et al. (2020)](doi:10.1186/s12991-020-00315-1)); serum albumin and depression ([Zhang et al. (2023)](doi:10.1186/s12888-023-04935-1)); serum uric acid levels and creatine phosphokinase ([Chen et al. (2023)](doi:10.1186/s12872-023-03333-5)); weight-adjusted-waist index and urinary incontinence ([Sun et al. (2024)](doi:10.1038/s41598-024-51216-2)); body shape index and prostate cancer ([Liu et al. (2024c)](doi:10.1007/s11255-023-03917-2)); sedentary behavior and urinary incontinence ([Di et al. (2024)](doi:10.1016/j.heliyon.2024.e27764)); triglyceride glucose body mass index and urinary incontinence ([Li et al. (2024)](doi:10.1186/s12944-024-02306-7)); dietary total energy intake and asthma ([Cao et al. (2024)](doi:10.1186/s40795-024-00938-7)); composite dietary antioxidant index and hyperlipidemia ([Zhao et al. (2024)](doi:10.1038/s41598-024-66922-0)); dietary zinc intake and asthma ([Cheng et al. (2024b)](doi:10.1016/j.waojou.2024.100900)); medical uninsurance and sociodemographic attributes ([Wahab et al. (2022)](doi:10.1097/MD.0000000000030539)); sleep health and albumin ([Li and Guo (2022)](doi:10.1186/s12889-022-13524-y)); visceral fat metabolic score and osteoarthritis ([Xue et al. (2024)](doi:10.1186/s12889-024-19722-0)); waist circumference and sex steroid hormones ([Zhu et al. (2024)](doi:10.1155/2024/4306797)); systemic immune-inflammation index and diabetes ([Nie et al. (2023)](doi:10.3389/fendo.2023.1245199)); systemic immune-inflammation index and hyperlipidemia ([Mahemuti et al. (2023)](doi:10.3390/nu15051177)); triglyceride glucose index and depression ([Ren et al. (2024)](doi:10.1097/MD.0000000000039258)); selenium levels and chronic kidney disease ([Pi et al. (2024)](doi:10.3389/fnut.2024.1396470)); composite dietary antioxidant index and cardiovascular disease ([Liu et al. (2023)](doi:10.3390/antiox12091740)); and systemic immune biomarkers and metabolic-associated fatty liver conditions ([Wang et al. (2024)](doi:10.3389/fnut.2024.1415484)). Each paper's exact exposure contrast, outcome, population, and cycles are in [results/associations.csv](results/associations.csv) and [plan/associations.json](plan/associations.json).

**Re-implementation.** Each association is one file, `code/associations/rowNNN.R`, built on a shared library (`code/lib/`): it builds the exposure, outcome, covariates, and study population from each cycle's NHANES files, lists every alternative tried on the paper's own cycles as a variant, and records every choice the paper left open with its reason. Models are fitted with the survey package ([Lumley (2004)](doi:10.18637/jss.v009.i08)) in R 4.6.1: weighted analyses by `svyglm` with Taylor-linearized variance over the masked strata and PSUs, the study population as a domain, pooled weights scaled by each cycle's share of the years, and t tests on the design's degrees of freedom (PSUs minus strata); unweighted analyses by `glm` with normal intervals. Ratio measures are analyzed on the log scale, and every interval is a 95% confidence interval. Unstated choices were settled by, in order, what the paper says elsewhere, NCHS's analytic guidelines, the convention of this literature, and, among choices still plausible, closeness to the published estimate. eGFR uses NCHS's calibration of 1999–2000 and 2005–2006 serum creatinine ([Selvin et al. (2007)](doi:10.1053/j.ajkd.2007.08.020)); the dietary inflammatory index uses its published global means and weights ([Shivappa et al. (2014)](doi:10.1017/S1368980013002115)); the composite dietary antioxidant index is the sum of six standardized intakes ([Wright et al. (2004)](doi:10.1093/aje/kwh173)).

What the replication tests is the published estimate, so the re-implementation follows the computation that produced it wherever the paper's own numbers reveal it, even where that departs from the paper's text: an unweighted analysis in a paper that says it weighted, a sample that complete-case coding restricted, a covariate the text names but the model left out. Two exceptions correct the computation instead: participants counted twice, by stacking the 2017–2018 files with the 2017–March 2020 files that contain them, are counted once, and an outcome coded in reverse is coded as the paper states it. Every such finding is recorded in the association's file as a departure, with its kind (the coding of a variable, the sample, the numbers reported, the survey weights or design, or the model), whether it bears on the headline estimate, and whether the re-implementation follows it; `code/run` counts the papers with each kind bearing on the headline. After registration, in answer to review, each departure bearing on the headline was also classified by how it is established: shown by the paper itself, where two of its parts disagree or arithmetic on its own printed numbers rules out what it describes; shown by recomputation, where the computation described does not reproduce the published numbers from the public files and another, identified, does; or unresolved, where recomputation rules out the computation described without identifying the one used. Under a written codebook, the lead agent and a second agent, blind to the first's labels, each coded every such departure from its recorded detail and its association's file, without returning to the papers; `data/departure_coding.csv` holds both codings, the lead agent settled their disagreements by rereading the details under the codebook, and the settled label is the departure's `evidence` in its file. One departure recorded as a model's in the registered files, where one section of a paper lists covariates that its model, as another section and its table give it, leaves out, is counted as a reporting departure here, since the paper describes its model both ways. Multiple imputation, whose draws cannot be reproduced, is replaced by complete cases.

**2021–2023.** The same code runs on the new cycle with rules fixed before registration. Each association runs in its harmonized version, which leaves out the covariates and exclusion steps 2021–2023 cannot build (`left_out` in each file) and is fitted on the paper's cycles too, so what leaving them out changes is measured. Cutpoints, standardization constants, and knots keep the values the paper used or, where it gave none, those computed on its cycles. The weight is the paper's, except that NCHS directs the phlebotomy weight for blood analytes and the first dietary recall's weight for the 30-day supplement questionnaire in this cycle; the cycle is analyzed alone, with 15 design degrees of freedom. A covariate with one value in the 2021–2023 sample would have been left out, and a model that could not be fitted would have counted as not replicated; neither happened.

**Outcomes.** An association **replicated** if its 2021–2023 estimate has the published sign and its p-value, corrected over the 40 associations by Benjamini and Hochberg's procedure ([Benjamini and Hochberg (1995)](doi:10.1111/j.2517-6161.1995.tb02031.x)), is below 0.05; the share replicated has an exact (Clopper-Pearson) interval. A test is **informative** if it had at least 0.80 power (two-sided, α = 0.05, noncentral t on the design's degrees of freedom) to detect the published effect given its 2021–2023 standard error. The **effect ratio** is the 2021–2023 estimate over the published one on the analysis scale, summarized by its median with a distribution-free interval. Each 2021–2023 estimate is tested against the published one (z on the analysis scale, the published standard error taken from its interval, corrected as above), and the effect ratio of each association is split into three factors: our estimate on the paper's cycles over the published one (what reproducing the paper gives), the harmonized estimate over ours (what leaving out what 2021–2023 lacks changes), and the 2021–2023 estimate over the harmonized one (what changed between cycles, tested by a second corrected z test). After registration, both tests were repeated with t distributions on the design degrees of freedom of the fits compared, and the tests' power was also computed at the Bonferroni level, 0.05/40: Benjamini and Hochberg's procedure rejects every p-value below that level and none above 0.05, so the power of the replication decision for a test lies between its power at those two levels. Intervals for shares and medians treat the associations as independent. All p-values are two-sided. In Table 1 and claims C10 to C12, the sleep estimate compares workday sleep of 5 hours or less with more than 7 and up to 8 hours, the diabetes estimate is per 100 units of the systemic immune-inflammation index, and the depression estimate is per unit of the triglyceride-glucose index.

**Data and code.** The bundle carries the 412 NHANES public-use files the analysis and the plan's dry run read, as CDC published them, compressed with xz, each checked against its recorded SHA-256 before it is read (`data/nhanes/sources.csv`; `code/fetch_data.R --check` compares them with CDC's current files). `code/run` re-runs everything from them in the image `env/Dockerfile` builds, in about 8 minutes on one core, and writes [results/R1.json](results/R1.json) (the aggregates), [results/associations.json](results/associations.json) (every fit and variant), and, through `code/report.R`, [results/associations.csv](results/associations.csv) and `results/R2.json` (the per-association values cited here).

**Ethics.** NHANES is conducted by the National Center for Health Statistics under protocols approved by its Research Ethics Review Board (#98-12, #2005-06, #2011-17, #2018-01, and #2021-05), with written informed consent. This study uses only the de-identified public-use files and attempts to identify no one. It reports what each published analysis computed, as the papers' own numbers show it, and makes no claim about how or why any paper was produced.

## Results

**The coding check.** On the cycles each paper analyzed, our code reproduced {{R1.reproduction.reproduced.k}} of the {{R1.reproduction.associations}} published estimates, its estimate falling inside the published interval, and our estimates were a median {{R1.reproduction.ratio_original.median}} times the published ones ({{R1.reproduction.ratio_original.ci.0}} to {{R1.reproduction.ratio_original.ci.1}}). Because open choices were settled partly by closeness to the published estimate (Methods), a reproduction shows that a computation consistent with the published numbers exists, not that it is the one the authors ran. In our own fits on those cycles, {{R1.reproduction.same_sign_p05.k}} estimates had the published sign and an unadjusted p-value below the threshold, fewer than published, partly because our intervals are design-based where several papers' were not and some of our complete-case samples are smaller. The estimate not reproduced is the association of dietary zinc with asthma ([Cheng et al. (2024b)](doi:10.1016/j.waojou.2024.100900)), whose estimates are the odds of having no asthma, as the paper's own tables show: our model of asthma, as the paper defines its outcome, gives {{R2.row257.original.estimate}} against the published {{R2.row257.published.estimate}}, and the reversed outcome gives {{R2.row257.variants.0.estimate}}.

**What the published analyses computed.** In {{R1.reproduction.departures.affecting_headline.k}} of the papers, the investigators recorded at least one departure bearing on the headline estimate, {{R1.reproduction.departures.headline_departures}} in all. Of these departures, {{R1.reproduction.departures.by_evidence.paper.k}} are shown by the paper itself, {{R1.reproduction.departures.by_evidence.data.k}} by recomputation that identifies another computation, and {{R1.reproduction.departures.by_evidence.unresolved.k}} are unresolved; the two independent codings agreed on {{R1.reproduction.departures.coding.agree.k}} of them (Cohen's kappa {{R1.reproduction.departures.coding.kappa}}), and {{R1.reproduction.departures.coding.settled}} disagreements were settled. By paper, {{R1.reproduction.departures.shown_by_paper.k}} papers have a departure the paper itself shows, {{R1.reproduction.departures.identified_by_data.k}} more have one that only recomputation shows, and {{R1.reproduction.departures.only_unresolved.k}} have only unresolved ones. By kind, as the investigators classified them, {{R1.reproduction.departures.by_kind.coding.k}} papers coded a variable otherwise than described, {{R1.reproduction.departures.by_kind.sample.k}} analyzed another sample, {{R1.reproduction.departures.by_kind.reporting.k}} reported a number or a description that disagrees with another part of the paper, {{R1.reproduction.departures.by_kind.weighting.k}} did not use the survey weights or design they describe, and {{R1.reproduction.departures.by_kind.model.k}} fitted a model other than the one stated. In {{R1.reproduction.departures.not_followed.k}} papers our analysis does not follow a departure, because it corrects it or because the paper's numbers rule out the stated computation without revealing the one used.

Some are large. The caffeine paper's headline, labeled the odds ratio for the highest against the lowest quartile, is an odds ratio per milligram within the highest quartile, as every cell of its table is and as its own obesity rates by quartile show ([Liu and Cui (2024)](doi:10.1371/journal.pone.0300566)); the contrast its text describes gives {{R2.row066.variants.5.estimate}}. The selenium paper's models are reproduced only without age, which its text lists as a covariate ([Pi et al. (2024)](doi:10.3389/fnut.2024.1396470)); with age its estimate is {{R2.row315.variants.3.estimate}} against the published {{R2.row315.published.estimate}}. The stroke paper on the systemic immune-inflammation index counted the participants of 2017–2018 twice ([Liu et al. (2024a)](doi:10.31083/J.RCM2504130)). Two published intervals are much narrower than a design-based analysis of the same data gives, because they ignore the survey design: the published interval is {{R2.row101.published.interval_ratio}} times the width of ours for blood manganese and liver stiffness ([Han et al. (2023)](doi:10.1186/s40001-022-00977-5)), and {{R2.row284.published.interval_ratio}} times for workday sleep and albumin ([Li and Guo (2022)](doi:10.1186/s12889-022-13524-y)). The gallstones paper's interval is {{R2.row040.published.interval_ratio}} times ours for another reason: its model, as the paper lists it, holds total and HDL cholesterol beside the log of their ratio, which they nearly determine, and its own interval shows it was not fitted that way, though which model was fitted is unresolved ([Cheng et al. (2024a)](doi:10.1186/s12944-024-02262-2)). Many other departures concern how a covariate was coded or which participants entered. Each departure is described, with its evidence and how it is established, in its association's file under `code/associations/`, and counted in [results/associations.csv](results/associations.csv).

**Replication.** In 2021–2023, {{R1.replication.replicated.k}} of the {{R1.replication.associations}} associations replicated (a share of {{R1.replication.replicated.share}}, {{R1.replication.replicated.ci.0}} to {{R1.replication.replicated.ci.1}}). One survey cycle holds fewer participants than the several most papers pooled, so only {{R1.replication.informative}} tests had the planned power, at the nominal level, to detect the published effect; of these, {{R1.replication.replicated_informative.k}} replicated ({{R1.replication.replicated_informative.share}}, {{R1.replication.replicated_informative.ci.0}} to {{R1.replication.replicated_informative.ci.1}}). Power at the nominal level bounds the power of the corrected decision from above; at the Bonferroni level, which bounds it from below, {{R1.replication.informative_bonferroni}} tests had the planned power, and {{R1.replication.replicated_informative_bonferroni.k}} of them replicated. The other tests had less power, so their failures to replicate are weak evidence either way. Effects in 2021–2023 were a median {{R1.replication.ratio.median}} times the published ones ({{R1.replication.ratio.ci.0}} to {{R1.replication.ratio.ci.1}}; quartiles {{R1.replication.ratio.quartiles.0}} and {{R1.replication.ratio.quartiles.1}}), and {{R1.replication.ratio_informative.median}} times ({{R1.replication.ratio_informative.ci.0}} to {{R1.replication.ratio_informative.ci.1}}) among the informative tests. {{R1.replication.in_published_ci.k}} estimates fell inside the published confidence interval, {{R1.replication.same_sign_p05.k}} had the published sign with an unadjusted p-value below the threshold, and {{R1.replication.reversed.k}} had the opposite sign with a significant corrected p-value. The count is the same among the reproduced associations ({{R1.replication.replicated_reproduced.k}} of {{R1.replication.replicated_reproduced.n}} replicated) and with each paper's own weight in place of the subsample weights NCHS directs for 2021–2023 ({{R1.replication.replicated_paper_weight.k}}). Every model fitted in 2021–2023, with no covariate left out for want of variation.

**Where 2021–2023 differs from the published estimates.** {{R1.replication.differs_from_published.k}} estimates differ significantly from the published ones after correction (Table 1), and {{R1.replication.differs_from_published_t.k}} with t tests on the fits' design degrees of freedom. Against our own harmonized estimates on the papers' cycles, {{R1.replication.heterogeneous.k}} differ significantly ({{R1.replication.heterogeneous_t.k}} with t tests), with a median ratio of {{R1.replication.ratio_own.median}} ({{R1.replication.ratio_own.ci.0}} to {{R1.replication.ratio_own.ci.1}}): a single cycle cannot tell most changes between cycles from sampling error, and finding none significant does not show the estimates equal.

**Table 1.** The associations whose 2021–2023 estimate differs significantly from the published one: the published estimate, ours on the paper's cycles (its own analysis), the harmonized one there, the 2021–2023 estimate with its confidence interval, the corrected p-value of the difference by z test and by t test, and the three factors of the effect ratio.

| Association | Published | Paper's cycles | Harmonized | 2021–2023 | Corrected p, z | Corrected p, t | Reproduction factor | Harmonization factor | Between cycles |
|---|---|---|---|---|---|---|---|---|---|
| Workday sleep and serum albumin, β (g/L) ([Li and Guo (2022)](doi:10.1186/s12889-022-13524-y)) | {{R2.row284.published.estimate}} | {{R2.row284.original.estimate}} | {{R2.row284.harmonized.estimate}} | {{R2.row284.replication.estimate}} ({{R2.row284.replication.low}} to {{R2.row284.replication.high}}) | {{R2.row284.difference_q}} | {{R2.row284.difference_q_t}} | {{R2.row284.ratio.original}} | {{R2.row284.ratio.harmonized}} | {{R2.row284.ratio.between_cycles}} |
| Immune-inflammation index and diabetes, OR ([Nie et al. (2023)](doi:10.3389/fendo.2023.1245199)) | {{R2.row303.published.estimate}} | {{R2.row303.original.estimate}} | {{R2.row303.harmonized.estimate}} | {{R2.row303.replication.estimate}} ({{R2.row303.replication.low}} to {{R2.row303.replication.high}}) | {{R2.row303.difference_q}} | {{R2.row303.difference_q_t}} | {{R2.row303.ratio.original}} | {{R2.row303.ratio.harmonized}} | {{R2.row303.ratio.between_cycles}} |
| Triglyceride-glucose index and depression, OR ([Ren et al. (2024)](doi:10.1097/MD.0000000000039258)) | {{R2.row311.published.estimate}} | {{R2.row311.original.estimate}} | {{R2.row311.harmonized.estimate}} | {{R2.row311.replication.estimate}} ({{R2.row311.replication.low}} to {{R2.row311.replication.high}}) | {{R2.row311.difference_q}} | {{R2.row311.difference_q_t}} | {{R2.row311.ratio.original}} | {{R2.row311.ratio.harmonized}} | {{R2.row311.ratio.between_cycles}} |

In all three, our analysis reproduces the published estimate on the paper's own cycles, and the difference arose between cycles. In the albumin paper short sleepers had lower albumin, and in 2021–2023 the estimate points the same way, closer to zero, with an interval that includes zero. The paper pooled 2015–2016 with 2017–2018, whose analyzer reads albumin lower, and short sleepers are more common in 2017–2018; with a cycle term its own model gives {{R2.row284.variants.9.estimate}} g/L, and in 2021–2023, analyzed alone, the estimate is {{R2.row284.replication.estimate}} g/L. Its intervals also ignore the survey design, and its weights are not the ones its text names. The diabetes paper analyzed unweighted data while saying it weighted them, and in 2021–2023 the estimate lies below one, though not significantly so after correction. The depression paper's model, as its table lists it and as reproduced, leaves out glucose and lipids, which another of its sections says were all included, and its 2021–2023 sample is {{R2.row311.replication.n}} people against {{R2.row311.original.n}} on its cycles, so the estimate below one there rests on few cases, and its difference from the published estimate is not significant by the t test (corrected p {{R2.row311.difference_q_t}}). The full table, with every association's estimates, tests, factors, and departures, is [results/associations.csv](results/associations.csv).

## Discussion

Most of the sampled associations could be neither confirmed nor refuted by one survey cycle, so the overall replication rate, {{R1.replication.replicated.k}} of {{R1.replication.associations}}, mostly measures power. Among the {{R1.replication.informative}} tests powered at the nominal level to detect the published effect, {{R1.replication.replicated_informative.k}} replicated, a share with a confidence interval from {{R1.replication.replicated_informative.ci.0}} to {{R1.replication.replicated_informative.ci.1}}. Those tests were chosen by their power against the published effect, which favors associations whose published effect is large relative to its uncertainty, so their rate need not carry over to the rest.

Replication projects that collected new data, most of them with high power, report lower rates and more shrinkage than the informative tests here. In psychology, 36% of 100 replications were statistically significant, against 97% of the originals, and their effects were half the original ones ([Open Science Collaboration (2015)](doi:10.1126/science.aac4716)). In experimental economics, 11 of 18 replications found a significant effect in the original direction, at 66% of the original effect on average ([Camerer et al. (2016)](doi:10.1126/science.aaf0918)), and for social science experiments published in Nature and Science, 13 of 21 did, at about half ([Camerer et al. (2018)](doi:10.1038/s41562-018-0399-z)). In preclinical cancer biology, the median replication effect of positive effects was 85% smaller than the median original one ([Errington et al. (2021)](doi:10.7554/eLife.71601)). Here, effects in 2021–2023 were a median {{R1.replication.ratio_informative.median}} times the published ones among the informative tests and {{R1.replication.ratio.median}} times over all, with intervals (over all, {{R1.replication.ratio.ci.0}} to {{R1.replication.ratio.ci.1}}) that include both no shrinkage and the halving those projects found. Those projects ran new studies with new teams; this study re-estimated observational associations in a new probability sample of the same population, with the same questions and measurements as far as the cycle allows and with the original analysis, so less could change between original and replication. Effects published because they crossed a significance threshold in underpowered studies are expected to be inflated, and flexible analyses with selective reporting can inflate them further ([Ioannidis (2008)](doi:10.1097/EDE.0b013e31818131e7)).

The coding check separates two things that replication projects usually see together: whether a published number can be recomputed, and whether the paper says how. Our code recomputed {{R1.reproduction.reproduced.k}} of the published estimates from the public files, at a median {{R1.reproduction.ratio_original.median}} times the published value, yet in {{R1.reproduction.departures.shown_by_paper.k}} papers the paper itself shows that the computation behind the number is not the one it describes, and in {{R1.reproduction.departures.identified_by_data.k}} more only another computation reproduced the number. In psychology, [Artner et al. (2021)](doi:10.1037/met0000365) reproduced 163 of 232 key statistical claims from the authors' raw data, 18 of them only by departing from the article's analytical description, often through cumbersome trial and error, and [Hardwicke et al. (2018)](doi:10.1098/rsos.180448) reproduced every target value in 22 of 35 articles with reusable data, 11 of them only with the authors' help, naming unclear analysis specification among the obstacles. NHANES differs in that every reader has the same data. Here each computation was found from the paper and the public files alone, by trying alternatives until one gave the published numbers. Such a match identifies a computation consistent with those numbers, not necessarily the one the authors ran, which is why the departures that the papers themselves show are counted apart; either way, a reader who takes the methods section at face value would not see the difference.

Ignoring a survey's design is a known error in secondary analyses. Among 145 analyses of a federal workforce survey, 55% used the sampling weights and 8% accounted for the complex sample in estimating variances, an earlier review of 100 articles on national health surveys suggested that such errors may be extremely prevalent, and ignoring the design can bias estimates and give standard errors that do not reflect the sample ([West et al. (2016)](doi:10.1371/journal.pone.0158120)). The weighting departures recorded here are narrower and harder to see: in {{R1.reproduction.departures.by_kind.weighting.k}} papers the text says the analysis used the weights or the design, and the published numbers are reproduced only without them.

[Suchak et al. (2025)](doi:10.1371/journal.pbio.3003152) documented these papers and raised two concerns about them: single-factor designs that skip false discovery correction, and selective use of survey cycles or subsets, which they call suggestive of data dredging. What this study adds is a test in data the papers could not have seen and a check of what their analyses computed. A cycle the papers could not have analyzed tests the second concern directly, since an association produced by choosing a date range or a subset has no reason to reappear in new data. Most of the informative tests found the association again, at a median {{R1.replication.ratio_informative.median}} times its published size, though that share rests on few tests; the other associations remain untested. A replication cannot answer the first concern: an association that reappears says nothing about how many others were tried, or about the factors a single-factor design leaves out. Nor is an estimate fixed by the data alone: [Patel et al. (2015)](doi:10.1016/j.jclinepi.2015.05.029) fitted 8,192 sets of adjustments to the association of each of 417 NHANES variables with mortality, and for 31% of the variables the direction differed between the 1st and 99th percentiles of those analyses. Several departures recorded here, such as a covariate the text lists but the model leaves out, are choices of that kind, made and not reported. Testing many factors at once, correcting for multiple comparisons, and validating findings in other NHANES cohorts, as an environment-wide association study does ([Patel et al. (2010)](doi:10.1371/journal.pone.0010746)), addresses false discovery and selective use of cycles alike. None of this bears on how or why any paper was written, about which this study claims nothing.

The three estimates that differ from their published values each have explanations other than an error in the original paper. In all three, our analysis reproduced the published estimate on the paper's own cycles, and leaving out what 2021–2023 lacks changed it little (the first two factors in Table 1), so the difference arose between cycles: from sampling error, from a change in the population or in how the cycle measured it, or from an original estimate that was inflated or confounded. For albumin, the laboratory analyzer changed between the two cycles the paper pooled, and short sleep was more common in the second, so the published estimate mixes the association with a measurement change; applying NCHS's comparison of the two analyzers, which the association's file records, to the paper's cycles would show how much of the estimate the change produced. For the immune-inflammation index and diabetes, the paper's unweighted analysis describes its sample rather than the population, so a design-based estimate in a new cycle need not agree with it even if nothing changed; a weighted analysis of the paper's own cycles would separate the two. The depression estimate rests on {{R2.row311.replication.n}} people, and its difference is not significant by the t test, so a later cycle may well find the published sign again. Only further cycles can tell sampling error from real change.

How unusual these departures are is not known. The papers were drawn from a literature already singled out as formulaic, and no comparable audit counts how often other NHANES papers, or observational studies in general, report numbers that come from a computation other than the one they describe. The same re-implementation applied to a random sample of other NHANES papers would tell whether the {{R1.reproduction.departures.affecting_headline.k}} papers here reflect this literature or the field. Until then, the departures speak to these papers, and the replication rate to these associations.

For anyone who relies on these papers, the results point three ways. Their estimates are the outputs of computations that may differ from the ones described, so anyone who cites, pools, or builds on one should reproduce it from the public files before relying on its methods, which every reader of an NHANES paper can do. Their effects are probably inflated: a meta-analysis or guideline that draws on them should expect later estimates to be smaller, by a median factor near {{R1.replication.ratio.median}} here, and weigh them accordingly. And an association this cycle could not test is untested rather than refuted: this study neither establishes nor rules out most of them.

For journals and reviewers, the departures in {{R1.reproduction.departures.shown_by_paper.k}} papers could be seen in the paper alone, in a table that contradicts the methods or a number that cannot follow from the computation described, yet all were published in peer-reviewed journals. Most of the others needed only the public files, which reviewers have as well. Asking for the analysis code, and checking that the stated weights, design, and sample give the reported numbers, is the closer attention to analytic methods that [West et al. (2016)](doi:10.1371/journal.pone.0158120) call for, and with public data it is a check that can be made before publication rather than after. The re-implementations here were done by AI agents from each paper and the public files, which suggests such checks can be made routine rather than left to a reader's spare time. For anyone pooling cycles, the albumin estimate (C10) shows a hazard: a change of laboratory method between cycles that coincides with a change in the exposure enters the pooled estimate unless the model has a term for the cycle. And the tests that lacked power stay open: within NHANES, only pooling 2021–2023 with later cycles, once they are released, can give them the power the plan required, and registering those tests before the data exist would keep them as clean as this one.

## Limitations

The main limitation is power: one two-year cycle has a fraction of the participants most papers pooled, so most tests lacked the planned power, and their failure to replicate says little. Power was registered at the nominal level, which overstates the power of the corrected decision; the count at the Bonferroni level bounds it from below. The effect ratios carry the most information, and their intervals are wide. The intervals for shares and medians treat the 40 associations as independent, but several share participants, cycles, exposures, or outcomes, so their errors may be correlated and the intervals are approximate. The tests against published estimates take the published standard error from a rounded interval and treat it as known. The 2021–2023 cycle also differs from its predecessors: response fell, the first dietary recall moved from an in-person interview to the telephone, blood pressure is measured by an oscillometric device, and the cycle came after the COVID-19 pandemic, any of which could change an association; a significant difference between cycles is therefore not by itself evidence that a published estimate was wrong. The harmonized versions leave out what 2021–2023 cannot build, which changed some estimates on the original cycles (the harmonization factor in `results/associations.csv`).

Following each paper's computation, rather than its text, was a judgment made from its own numbers, and open choices were settled partly by closeness to the published estimate, so a reproduced estimate is compatible with the published one rather than proof of the authors' computation; settling them before 2021–2023 was seen keeps that choice from favoring the replication. The departures recorded are those the coding check found while reproducing the headline estimate; papers' other tables were checked unevenly, so departures that do not bear on the headline are listed but not counted. Departures identified only by recomputation rest on a reconstruction matching the published numbers, and we did not measure how far each departure moves the headline estimate; some move it little. Their kinds are the investigators' classification, and how each is established was classified after registration, in answer to review, by the lead agent and a second agent of the same model family working blind to the first, not by an outside adjudicator. Where a paper's computation could not be identified, as for the sample of the gallstones paper or the antioxidant index of the cardiovascular paper, our estimate may differ from what its authors computed for reasons we cannot see. Multiple imputation was replaced by complete cases. The sample covers papers eligible under the plan's rules, and excludes papers without full text in Europe PMC (53 of the 206 judged), which may favor open-access journals; 85 papers were judged ineligible from Table A's labels and abstract alone. Extraction, coding, and the classification of departures were done by AI agents of one model family, reviewed by its lead agent but not by people, and every judgment is recorded, with its reason, in the files a reader can check. The study tests associations, not causal effects, and none of the associations, replicated or not, should be read as causal.

## Provenance

The study was planned, run, and written by AI agents of the Claude family (claude-opus-5-5). A lead agent searched prior work, wrote the plan, the eligibility and headline rules, and the shared R library, coordinated agents working under written briefs who screened papers, extracted each paper with verbatim quotes, coded each association, and recorded its departures, reviewed and corrected their work, registered the plan, ran the analysis, and wrote the claims and this paper. For this corrected version, the lead agent took up the first version's reviews, a second agent coded how each departure is established without seeing the lead agent's coding, and another agent searched and checked the earlier work the Discussion cites. The tools were R 4.6.1 with the survey, jsonlite, and foreign packages in a pinned Docker image, Python's standard library for the plan's tables, and Europe PMC's and Crossref's public APIs. The data are the NHANES public-use files, Table A of S1 Data of [Suchak et al. (2025)](doi:10.1371/journal.pbio.3003152) under CC BY 4.0, and the sampled papers' full texts, which the bundle does not carry. A person asked for studies outside mathematics; the lead agent proposed this one among others, and the person chose it. The person saw a summary of the plan before registration and approved it unchanged, asked that the report highlight where the 2021–2023 results differ from the published ones and what those analyses computed, and, on reading a summary of the findings, said the bottom line read too softly, after which the agent rewrote the first version's title and Summary. After the first version's reviews, the person asked whether its title was accurate and, when the agent judged that it overstated the findings, asked for this correction of the title and of what the reviews found. Every other proposal, from the plan's rules to the record of departures and this version's changes, was the agents', and the person accepted them. No person reviewed the extractions, the code, the results, or this paper. `provenance.json` says the same.
