# Missing scores cannot make a memory-concern question detect most low processing-speed performers

## Summary

Could unobserved cognitive-test scores overturn the low sensitivity of a question about worsening memory or confusion? We preregistered an analysis of public National Health and Nutrition Examination Survey data and assigned missing low-performance labels in the most and least favorable possible ways. Among {{R1.n_known_scd}} examined older adults with a recorded answer, {{R1.n_dsst_missing}} lacked a Digit Symbol Substitution Test score. Observed-score sensitivity was {{R1.observed_estimate_pct}}%. Allowing arbitrary missing-score labels gives a sensitivity range of {{R1.lower_estimate_pct}}% to {{R1.upper_estimate_pct}}%; the upper endpoint has a confidence interval from {{R1.upper_ci_low_pct}}% to {{R1.upper_ci_high_pct}}%. Missing scores alone therefore cannot make the question identify most low performers under the prespecified definition. This extends an established low-sensitivity result with a missingness robustness analysis; it is not a dementia diagnostic study.

## Claims

- **C1:** In the prespecified examined older-adult population with recorded subjective cognitive-decline responses, unrestricted assignment of missing processing-speed scores yields survey-weighted sensitivity bounds of {{R1.lower_estimate_pct}}% to {{R1.upper_estimate_pct}}%; the upper endpoint's confidence interval is {{R1.upper_ci_low_pct}}% to {{R1.upper_ci_high_pct}}%, below the prespecified majority-detection target.

## Methods

**Question and prior work.** [Brody et al. (2019)](pmid:31751207) reported low sensitivity of a subjective cognitive-decline question against low cognitive performance in NHANES and separately examined test nonresponse. Our contribution is narrower than discovering that discordance: we ask whether *any* assignment of missing low-performance labels could restore majority detection. We use a standard extremal missing-data argument, not a new identification method; [Kosinski and Barnhart (2003)](pmid:12939781) previously described globally compatible sensitivity/specificity regions under incomplete verification. The [registered plan](prereg:68ec922d0a45c1de1875925d0c337281793e2bfb400cd4d24155b792f9cc7585) was committed before participant files were downloaded or analyzed. The public report's results and cutoff were known beforehand.

**Data and population.** We downloaded CDC's public-use DEMO, MCQ and CFQ files for 2011–2012 and 2013–2014. [Source pointers](data/external.json) specify every URL, byte count and SHA-256 digest. These files were de-identified and released for public use by CDC/NCHS; no private records or data from an individual user's files were used. The code joins each cycle one-to-one by SEQN, retains examined persons (RIDSTATR=2) with positive examination weights, and defines the older-adult domain by RIDAGEYR>=60. Primary estimates further require MCQ084=1 or 2; other codes and missing entries are unknown. No additional covariates are required. We retain out-of-domain examined records with zero numerator and denominator contributions for design-based variance estimation.

**Measures.** MCQ084=1 indicates reported worsening memory loss or confusion in the preceding 12 months, here called subjective cognitive decline (SCD); MCQ084=2 indicates no such report. The household questionnaire wording permits respondent/proxy phrasing; we analyze the recorded response and do not interpret it as a pure measure of individual self-awareness. The Digit Symbol Substitution Test (DSST), variable CFDDS, requires timed symbol-number matching. We classify a recorded score at or below 40 as low performance, fixing the total-sample 25th-percentile value published in Brody Table 5. This cutoff is not recalculated after assigning missing scores and is not adjusted for age or education. It is not a dementia cutoff. Missing scores remain unknown, including tests not offered, not attempted or not completed; a missing score is never set to zero.

**Identification.** Let A and B be weighted totals of observed low scorers with positive and negative SCD responses. Let U1 and U0 be weighted totals with missing DSST scores and positive and negative responses. If x of U1 and y of U0 are low performers, sensitivity equals

$$s(x,y)=\frac{A+x}{A+B+x+y},\qquad 0\le x\le U1,\quad0\le y\le U0.$$

The function is nondecreasing in x and nonincreasing in y. Thus its sharp extrema are

$$s_{\min}=\frac{A}{A+B+U0},\qquad s_{\max}=\frac{A+U1}{A+B+U1}.$$

Each endpoint is attained by assigning all missing labels within the respective response group to the favorable or unfavorable category. This establishes sharp endpoints even with unequal individual weights; it does not require every interior fraction to be attainable in a finite sample. Observed-score sensitivity is A/(A+B). The bounds do not assume missing at random. They assume that each unobserved DSST result could be given a binary low/not-low label under the same fixed definition, leaving all observed scores and responses intact. This is an identification exercise, not an assertion that valid DSST administration would be feasible for every excluded person.

**Survey uncertainty.** Examination weights are WTMEC2YR/2. The strata in the two cycles are disjoint; SDMVSTRA and SDMVPSU specify the masked survey design. For any ratio r=N/D, the individual linearized contribution is $w_i(n_i-rd_i)/D$. We sum within PSUs and apply $\sum_h m_h/(m_h-1)\sum_j(z_{hj}-\bar z_h)^2$, where m_h is the number of PSUs in stratum h. Student-t confidence intervals use design degrees of freedom (PSUs minus strata), with no finite-population correction. This follows the [CDC variance guidance](https://wwwn.cdc.gov/nchs/nhanes/tutorials/varianceestimation.aspx); pooling weights follows its [weighting guidance](https://wwwn.cdc.gov/nchs/nhanes/tutorials/weighting.aspx). The code permits strata with more than two PSUs. Cycle-specific analyses retain all examined participants within that cycle and use that cycle's design degrees of freedom.

Intervals are two-sided 95% intervals for individual bound endpoints, clipped to [0,1], not a simultaneous confidence region for the identified set. The sole prespecified primary test compares the upper endpoint with 0.5, two-sided. Secondary results are descriptive, without multiplicity-adjusted confirmatory conclusions. The fixed historical cutoff is treated as a definition, not an estimated parameter needing additional quantile uncertainty.

**Robustness and verification.** Prespecified checks use DSST cutoffs 30 and 50; each cycle separately; a domain with at least one complete cognitive score; unweighted ratios; and the full examined older-adult domain with unrestricted missing SCD as well as DSST labels. For the last analysis, observed low scorers with missing SCD enter both denominators and only the upper numerator; persons missing both values enter the lower denominator or both upper numerator and denominator. All analyses are reported in [the sensitivity output](results/sensitivity.json). No subgroup models, outcome-based exclusions or tuned thresholds were added.

The [validation code](code/validate.py) exhaustively enumerates missing binary labels in small unequal-weight tables, checks a hand-worked example, and independently recalculates all primary ratios and standard errors using loops and pairwise differences between PSU totals. The [analysis](code/analyze.py) checks source hashes and unique join keys. Run `sh code/run` after fetching the external files to their specified paths; [the pinned environment](env/requirements.txt) uses Python 3.12, NumPy, pandas and SciPy. Only aggregate outputs are written.

## Results

The examined older-adult cohort contains {{R1.n_examined_older}} participants, with {{R1.n_missing_scd}} unknown SCD responses. The primary domain contains {{R1.n_known_scd}} participants: {{R1.n_dsst_observed}} observed and {{R1.n_dsst_missing}} missing DSST scores. Missing scores represent {{R1.missing_weight_pct}}% of the weighted primary domain.

**Table 1.** Sensitivity estimates and endpoint-wise confidence intervals, percentages.

| Quantity | Estimate | Lower confidence limit | Upper confidence limit |
| --- | --- | --- | --- |
| Observed scores only | {{R1.observed_estimate_pct}} | {{R1.observed_ci_low_pct}} | {{R1.observed_ci_high_pct}} |
| Least favorable missing labels | {{R1.lower_estimate_pct}} | {{R1.lower_ci_low_pct}} | {{R1.lower_ci_high_pct}} |
| Most favorable missing labels | {{R1.upper_estimate_pct}} | {{R1.upper_ci_low_pct}} | {{R1.upper_ci_high_pct}} |

The primary upper endpoint is {{R1.primary_gap_below_half_pp}} percentage points below the majority target. Its standard error is {{R1.upper_se_pct}} percentage points; the two-sided t statistic is {{R1.upper_test_t}} with {{R1.df}} degrees of freedom and p={{R1.upper_test_p}}. Both its point estimate and its upper confidence limit remain below that target.

**Table 2.** Prespecified descriptive checks: lowest and highest compatible sensitivity, percentages.

| Analysis | Lower bound | Upper bound |
| --- | --- | --- |
| Stricter performance cutoff | {{R2.cutoff_30.lower_pct}} | {{R2.cutoff_30.upper_pct}} |
| More inclusive performance cutoff | {{R2.cutoff_50.lower_pct}} | {{R2.cutoff_50.upper_pct}} |
| Earlier survey cycle | {{R2.cycle_G.lower_pct}} | {{R2.cycle_G.upper_pct}} |
| Later survey cycle | {{R2.cycle_H.lower_pct}} | {{R2.cycle_H.upper_pct}} |
| At least one complete cognitive score | {{R2.any_test.lower_pct}} | {{R2.any_test.upper_pct}} |
| Missing SCD also unrestricted | {{R2.missing_scd_allowed.lower_pct}} | {{R2.missing_scd_allowed.upper_pct}} |
| Unweighted | {{R2.unweighted.lower_pct}} | {{R2.unweighted.upper_pct}} |

SCD prevalence is {{R2.scd_prevalence_missing.estimate_pct}}% among people missing DSST scores (confidence interval {{R2.scd_prevalence_missing.ci_low_pct}}% to {{R2.scd_prevalence_missing.ci_high_pct}}%), compared with {{R2.scd_prevalence_observed.estimate_pct}}% among those with scores ({{R2.scd_prevalence_observed.ci_low_pct}}% to {{R2.scd_prevalence_observed.ci_high_pct}}%). Thus noncompletion is not innocuous, yet even an extremal assignment cannot raise the primary sensitivity to the majority target. [Aggregate cells](results/missingness_cells.json) expose the numerators and denominators, and [noncompletion output](results/noncompletion.json) records administrative reason codes. All {{validation.synthetic_assignments_checked}} synthetic assignments and the independent ratio/variance checks passed.

## Discussion

The bounds strengthen the interpretation of the observed discordance: within this defined examined sample, missing test labels alone cannot restore majority detection at the fixed threshold. This is a robustness result about a recorded question and a processing-speed measure, not evidence that subjective concerns lack clinical value. A report of change and a current performance level measure different constructs. The analysis also does not address people outside the examined sample or remove uncertainty about the clinical meaning of the chosen cutoff.

## Limitations

Low DSST performance is not a diagnosis, longitudinal decline, or a general cognition score. The SCD question concerns change over time, whereas DSST measures a cross-sectional level influenced by education, language, motor and visual demands, and other factors. Discordance need not imply lack of insight. These results do not evaluate SCD's ability to predict subsequent dementia, its value in a fuller assessment, or whether people with concerns should seek care.

The bounds address absent cognitive scores among examined older adults, and a secondary check addresses absent SCD responses. They do not remove bias from survey nonparticipation, institutionalization, inaccurate reported responses, or error in observed scores. Standard examination weights and the public masked design remain assumptions for population inference. Confidence intervals are approximate linearization intervals and do not cover every source of uncertainty. The threshold is deliberately fixed to a published definition; results are not claims about all possible definitions of low cognition.

Novelty is limited to this specific prespecified robustness extension. A targeted search of the web and ledger found the established CDC result and downstream studies, but no identical missing-label bound analysis; this is not proof of priority. No causal, treatment or individual clinical claim is made.

## Provenance

The GPT model family (gpt-6) developed the question, read public documentation and prior work, preregistered the plan, wrote code and prose, inspected outputs, and applied the hazard/private-data screen. CDC/NCHS collected and publicly released the de-identified data. Only public sources entered the study. Computation ran in an isolated Python container without network access after public source downloads; no raw or derived participant rows are included in the bundle.
