# Exact CDC WONDER request specification

This is the query specification for the sciencejournal.ai naloxone DiD follow-up. It describes final mortality by **state of residence**, not state of occurrence, and calendar year. The input TSVs retain WONDER's own Query Parameters, citations and notes. The parser verifies those parameters before publishing.

## Why browser export

Official [NCHS public-use microdata omit state/county/city identifiers from 2005 onward](https://www.cdc.gov/nchs/nvss/dvs_data_release.htm), so they cannot supply this 2008-2021 state panel. The [NCHS bulk drug-poisoning table](https://data.cdc.gov/National-Center-for-Health-Statistics/NCHS-Drug-Poisoning-Mortality-by-State-United-Stat/xbxb-epbu) describes drug poisoning, not the requested MCOD opioid intersection and T40.4 subset. No matching bulk substitute is vendored. The exporter therefore uses the public browser form rather than the restricted vital-statistics API.

The official [MCOD dataset index](https://wonder.cdc.gov/mcd.html) separates the final bridged-race dataset ending in 2020 from the final single-race dataset starting in 2018. Use these final datasets, not provisional data or Underlying Cause of Death alone:

| Export | Dataset / request URL | Years selected | MCOD filter |
| --- | --- | --- | --- |
| `opioid_D77_2008_2020.tsv` | [Multiple Cause of Death, 1999-2020, D77](https://wonder.cdc.gov/mcd-icd10.html) | Every year 2008-2020 | Opioid list below |
| `t40_4_D77_2008_2020.tsv` | D77, same URL | Every year 2008-2020 | T40.4 only |
| `opioid_D157_2021_2021.tsv` | [Multiple Cause of Death, 2018-2024, Single Race, D157](https://wonder.cdc.gov/mcd-icd10-expanded.html) | 2021 only | Opioid list below |
| `t40_4_D157_2021_2021.tsv` | D157, same URL | 2021 only | T40.4 only |

The single-race dataset's available end year may advance; the selected year remains 2021. Default analysis coverage is 2008-2021. If `--start-year 1999` is used, select 1999-2020 for D77 and use filenames ending `_1999_2020.tsv`.

## Form settings for every request

1. Open the relevant dataset and click **I Agree** to the data-use restrictions.
2. **Organize table layout**: Group Results By **State**; And By **Year**; remaining three groupings **None**. Internal indexes are `B_1=D77.V9-level1`, `B_2=D77.V1-level1` (substitute `D157` for the single-race dataset). Do not group by cause of death, age, race, sex, county or month. Do not sum cause-specific mention rows.
3. **Measures**: Deaths (`M_1`), Population (`M_2`), Crude Rate (`M_3`), **Age Adjusted Rate** (`O_aar_enable`), its **95% Confidence Interval** (`O_aar_CI`) and **Standard Error** (`O_aar_SE`). Crude-rate CI/SE and Percent of Total are off. Optional crude CI/SE columns in manual exports are accepted and retained in raw inputs.
4. Open **Additional Rate Options**: Calculate Rates Per **100,000** (`O_rate_per=100000`). Standard Population **2000 U.S. Std. Population** (`O_aar_pop=0000`), rather than the separate rounded standard-million choice. **Use non-standard rates off**. **Archive Rates off** where present. All standard-age/population filters unrestricted.
5. **Select location**: States/Counties of residence (`O_location=D77.V9` or `D157.V9`), **The United States / All**. Include all 50 states plus DC. No territories. Urbanization: all categories; use the default 2013 classification with no restriction. D157 occurrence-location fields remain inactive.
6. **Select demographics**: **All Ages** (Ten-Year Age Groups, `O_age=<dataset>.V5`, `V_<dataset>.V5=*All*`), **All Sexes**, **All Races**, **All Hispanic Origins**, including unknown categories in the all selections. D157 uses the **6 single race categories** mode, All Races. Do not select infant ages or restrict age strata. Age is not a grouping index.
7. **Select year and month**: Advanced Finder Options; enter each selected calendar year on its own line in the **Selected Items** box (`V_<dataset>.V1`). No month restriction. Year/Month footer must contain the complete selected years.
8. **Underlying cause of death**: choose **ICD-10 Codes** (`O_ucd=<dataset>.V2`). Use Advanced Finder Options and paste the expanded list below into **Selected Items** (`V_<dataset>.V2`). WONDER rejects the shorthand `X40-X44` in this text box, so every individual code is listed.
9. **Multiple cause of death**: choose **ICD-10 Codes** (`O_mcd=<dataset>.V13`). Paste the appropriate query's code list into the top **Select Records with any of these items** box (`V_<dataset>.V13`). Leave the lower **AND any of these items** box **empty** (`V_<dataset>.V13_AND`). Codes within the top box are OR alternatives, not an AND requirement.
10. **Other options**: **Show Totals off**, **Show Zero Values on**, **Show Suppressed Values on**, precision **1 decimal place**, timeout **10 minutes**. Keep all weekdays, autopsy statuses and places of death unrestricted. Leave automatic Export Results off so that the result table is inspected first. Click **Send** once.
11. Wait for the results table, including the age-adjusted-rate measure. Click **Export**, choose **TSV (Tab Separated Values)** and **Download**. Save under the corresponding filename above. Preserve the entire export, including the Query Parameters footer. The script converts these official TSVs to CSV.

Death counts and rates based on confidential small cells remain suppressed. Unreliable rates remain flagged as WONDER reports them. Missing or unreliable rates are not turned into zeros. Totals can be marked `Disabled` by WONDER because of suppression constraints, which the parser accepts. An incomplete set of state-years causes failure, with no placeholder rows added.

## Exact underlying-cause list for both definitions

This is `X40-X44`, `X60-X64`, `X85`, `Y10-Y14`, including unintentional, suicide, homicide and undetermined intent:

```text
X40
X41
X42
X43
X44
X60
X61
X62
X63
X64
X85
Y10
Y11
Y12
Y13
Y14
```

## Exact opioid multiple-cause list

```text
T40.0
T40.1
T40.2
T40.3
T40.4
T40.6
```

An opioid case is **one death with the underlying cause in the first list AND at least one mention from this opioid list**. Each selected certificate contributes once to the state-year count, even if it mentions several opioids. T40.5 is not selected. The definition follows the requested case definition, also documented in [NCHS Data Brief 428](https://www.cdc.gov/nchs/products/databriefs/db428.htm).

## Exact T40.4 multiple-cause list

```text
T40.4
```

Keep the underlying-cause list identical. T40.4 identifies synthetic opioids other than methadone, which the WONDER finder labels “Other synthetic narcotics.” This is a subset of the opioid definition, not a mutually exclusive cause category. A death may mention T40.4 and another opioid. The share is **T40.4-involved count / opioid-involved count** within each state-year. It is not a ratio of age-adjusted rates and is unavailable where either count is suppressed/missing.

## Rates and population provenance

The selected AAR, lower/upper 95% confidence limits and SE come directly from WONDER. No age-stratum formula is implemented. [D77 documentation](https://wonder.cdc.gov/wonder/help/mcd.html) describes the rate options and population series; deaths with unknown age are included in counts and crude rates but excluded from AAR calculations. D77 uses the current default bridged-race estimates and census counts, including revised intercensal estimates for 2001-2009. The raw export's Rate Options field is “Default intercensal populations for years 2001-2009 (except Infant Age Groups).”

For 2021, [D157 population documentation](https://wonder.cdc.gov/wonder/help/mcd-expanded.html#Population) describes single-race Census Vintage 2021 July 1 resident estimates using a blended base. The 2021 denominator methodology differs from earlier years. Both queries use the same dataset, population and settings within each year. Inspect this transition when interpreting trends; the exporter records it without substituting another population. The generic default Rate Options text can still appear in D157's footer, even though only 2021 is selected.

## Expected input column map

| Official TSV/CSV column | Output |
| --- | --- |
| State | `state`, full state name |
| State Code | `state_fips`, two-character text |
| Year | `year` |
| Deaths, opioid query | `deaths`, `opioid_deaths` |
| Deaths, T40.4 query | `t40_4_deaths` |
| Population | `population` |
| Crude Rate | `crude_rate` |
| Age Adjusted Rate | `age_adjusted_rate` |
| Age Adjusted Rate Lower 95% Confidence Interval | `aa_rate_ci_lower` |
| Age Adjusted Rate Upper 95% Confidence Interval | `aa_rate_ci_upper` |
| Age Adjusted Rate Standard Error | `aa_rate_se` |
| Notes | `opioid_source_notes`, `t40_4_source_notes` |

The exact labels with “Per 100,000” on the two rate columns are also accepted. `Year Code` is permitted. The parser rejects unexpected grouping columns, missing required AAR/CI/SE columns, wrong query criteria, duplicates, row gaps, unrecognized marker values, incompatible populations and T40.4 counts exceeding opioid counts. It joins by state FIPS and year, not source row order. SHA-256 covers the complete raw inputs and all delivered snapshot files.
