# Recent warming acceleration tests depend on endpoints and the assumed noise model

## Summary

How sensitive is evidence for recent global warming acceleration to the final years included and the model of natural variability? A registered audit uses a frozen NASA global annual temperature series. A fixed-breakpoint model estimates a warming-rate increase of {{R1.display.primary_delta}} °C per decade, with a confidence interval from {{R1.display.primary_low}} to {{R1.display.primary_high}}. A search over possible breakpoints gives p = {{R1.display.endpoints.2025.search_p}} with fitted autocorrelation, but p = {{R1.display.endpoints.2022.search_p}} when the recent endpoint years are excluded. Stronger assumed autocorrelation also weakens the test. These are conditional statistical results about an unadjusted series, with limited simulated power for smaller slope changes.

## Claims

- **C1:** In the frozen NASA series, the fixed-breakpoint warming-rate increase is {{R1.display.primary_delta}} °C per decade, while the breakpoint-search p-value changes from {{R1.display.endpoints.2025.search_p}} with the full endpoint to {{R1.display.endpoints.2022.search_p}} with the earlier endpoint and {{R1.display.rho_sensitivity.rho_0p6}} under the stronger registered autocorrelation assumption.

## Methods

This is a conditional sensitivity audit, not a test of whether anthropogenic warming exists, a climate attribution analysis, a future projection, or a reproduction of all the methods in an earlier paper. Recent acceleration and its detectability already have a literature. [Beaulieu et al. (2024)](doi:10.1038/s43247-024-01711-1) examine breakpoint detection in global temperature records. [Foster and Rahmstorf (2026)](doi:10.1029/2025gl118804) analyze several records after estimating and removing El Niño/Southern Oscillation, volcanic and solar variability. Their adjusted analyses and the present unadjusted audit have different inputs and noise models. This audit offers a frozen, offline implementation of registered endpoint and noise sensitivities. No first-discovery claim is made.

The data are NASA GISTEMP version 4 global land-ocean annual anomalies relative to 1951–1980, obtained from the [NASA data service](https://data.giss.nasa.gov/gistemp/) during the session, with retrieval time, exact source URL and SHA-256 in [source provenance](data/source.json). The [raw CSV](data/gistemp.csv) is preserved without edits. [Lenssen et al. (2024)](doi:10.1029/2023jd040179) describe the observational uncertainty ensemble for this series; this audit uses the central estimate, without propagating that ensemble. NASA revises past entries as observations and methods change, so the results concern the frozen bytes. The analysis uses the J-D annual column for 1970–2025, with complete endpoints at 2025, 2024 and 2022. Every retained year must have twelve finite monthly entries and a finite annual entry. Partial 2026 observations are excluded. Unique years, consecutive coverage and agreement between the separately rounded annual and monthly entries are checked.

The [plan](plan/analysis-plan.json) was [registered](prereg:8d85f0f27d13f5159f7421a9d7ef42ba64655f1b889daffdc2ed444562a3c2d5) before fitting any model or inspecting downloaded outcome values. Published warming and acceleration findings, including the discussion of change near 2015, were already known. The file was downloaded before registration. This is a registration of analysis choices, without blinding to the literature. Unstated numerical implementation choices are listed in [deviations](deviations.json).

The primary model is a continuous linear spline fitted by ordinary least squares:

$$T_y=\alpha+\beta x_y+\delta h_y+\epsilon_y,\quad x_y=(y-1970)/10,\quad h_y=\max(0,y-2015)/10.$$

Here $T_y$ is the annual anomaly in °C, $y$ the calendar year, $\alpha$ the intercept, $\beta$ the earlier warming rate, $\delta$ the change in rate, and $\epsilon_y$ the residual. Rates are in °C per decade. The knot is fixed at 2015 from prior literature. The primary endpoint is 2025. A positive $\delta$ represents an increase in the fitted rate, rather than a continuously increasing second derivative. The primary two-sided test is $\delta=0$, with a 95% normal confidence interval. Its heteroskedasticity and autocorrelation consistent (HAC) covariance uses the Bartlett-weighted Newey-West estimator with lag three and the factor $n/(n-3)$, where $n$ is the number of annual observations. [Newey and West (1987)](doi:10.2307/1913610) describe this covariance estimator. The same fit is reported for HAC lags zero, one, three and five at all registered endpoints, together with the conventional independent-error t interval. These sensitivity tests are descriptive and unadjusted, without additional confirmatory claims.

A separate search compares a straight line with a continuous spline containing a single hinge at each integer year from 1985 through the earlier of 2015 and the endpoint minus ten. The statistic is the largest extra-sum-of-squares F statistic for adding a hinge, with residual degrees of freedom $n-3$. It tests departure from a straight line in either direction; it does not select only positive slope changes. Searching is accounted for by repeating the entire search in every simulated series. Under the null, residuals follow a stationary, zero-mean Gaussian first-order autoregression, abbreviated AR(1). The lag coefficient $\rho$ is estimated by regressing the null-model residual on its predecessor, without an intercept, and clipped to the range $[-0.95,0.95]$ if needed. Innovation variance is the mean squared recursion residual. The initial draw has the stationary variance, and each following draw uses the recursion. All regressions are refitted in each replicate through an equivalent residualized-hinge projection.

Each endpoint uses 10,000 simulations, NumPy PCG64 and seed 20261006. The reported Monte Carlo p-value is $(b+1)/(B+1)$, where $b$ is the exceedance count and $B$ the simulation count. A conditional binomial Monte Carlo standard error and Wilson interval for the exceedance probability are supplied. They describe simulation error, not uncertainty in the estimated noise parameters. Independent-error searches use $\rho=0$ and the null residual variance with $n-2$ degrees of freedom. At the 2025 endpoint, additional registered calibrations fix $\rho$ at 0.2, 0.4 and 0.6, reestimating innovation variance from the same null residuals. These are assumptions for sensitivity analysis, not competing estimates of the actual residual dependence.

The conditional power experiment uses the fitted 2025 null noise parameters, a true knot at 2015, and slope increases of 0.1, 0.2 and 0.3 °C per decade. Each setting has 5,000 simulated series. A replicate rejects if its maximum search statistic exceeds the empirical 95th percentile of the corresponding null simulations. This uses the full search rejection rule and reports binomial Monte Carlo uncertainty. The alternatives were specified in advance. Power is conditional on this Gaussian AR(1) model and the estimated scale; it is not power under every plausible climate process. There is no claim that a non-significant test rules out acceleration.

The [runner](code/run) executes the analysis offline in the [pinned environment](env/Dockerfile). A separately implemented scalar sequence of least-squares fits validates the vectorized search on the observed series and ten null replicates per calibration. The manually implemented HAC covariance is cross-checked against statsmodels at every registered lag and endpoint. All comparisons must pass. The code generates the tables, figures and every result in [the complete result record](results/R1.json). JSON floating-point summaries retain ten significant digits for portability, while the paper uses display rounding appropriate to the statistical uncertainty. These are author-run cross-checks, not independent ledger verification.

## Results

The primary fit uses {{R1.primary.n}} annual observations. Its earlier rate is {{R1.display.primary_before}} °C per decade and its later rate {{R1.display.primary_after}} °C per decade. Their difference is {{R1.display.primary_delta}} °C per decade, with standard error {{R1.display.primary_se}} and a {{R1.display.confidence_percent}}% normal HAC confidence interval from {{R1.display.primary_low}} to {{R1.display.primary_high}} °C per decade. The two-sided nominal p-value is {{R1.display.primary_p}}. This interval is conditional on the fixed knot and the specified covariance estimator.

**Table 1.** Endpoint sensitivity of the fixed-knot fit and the search-adjusted tests. Confidence intervals use the registered HAC estimator; the search tests use fitted autoregressive noise or independent errors.

| Endpoint | Annual observations | Rate increase, °C/decade | Confidence interval, °C/decade | Fixed-knot p | Fitted null lag coefficient | Autoregressive search p | Monte Carlo SE of search p | Independent-error search p |
|---|---:|---:|---|---:|---:|---:|---:|---:|
| 2025 | {{R1.fixed_knot_sensitivity.2025.3.n}} | {{R1.display.endpoints.2025.delta}} | {{R1.display.endpoints.2025.low}} to {{R1.display.endpoints.2025.high}} | {{R1.display.endpoints.2025.fixed_p}} | {{R1.display.endpoints.2025.rho}} | {{R1.display.endpoints.2025.search_p}} | {{R1.display.endpoints.2025.mc_se}} | {{R1.display.endpoints.2025.iid_search_p}} |
| 2024 | {{R1.fixed_knot_sensitivity.2024.3.n}} | {{R1.display.endpoints.2024.delta}} | {{R1.display.endpoints.2024.low}} to {{R1.display.endpoints.2024.high}} | {{R1.display.endpoints.2024.fixed_p}} | {{R1.display.endpoints.2024.rho}} | {{R1.display.endpoints.2024.search_p}} | {{R1.display.endpoints.2024.mc_se}} | {{R1.display.endpoints.2024.iid_search_p}} |
| 2022 | {{R1.fixed_knot_sensitivity.2022.3.n}} | {{R1.display.endpoints.2022.delta}} | {{R1.display.endpoints.2022.low}} to {{R1.display.endpoints.2022.high}} | {{R1.display.endpoints.2022.fixed_p}} | {{R1.display.endpoints.2022.rho}} | {{R1.display.endpoints.2022.search_p}} | {{R1.display.endpoints.2022.mc_se}} | {{R1.display.endpoints.2022.iid_search_p}} |

The selected search knot is {{R1.bootstrap.2025.AR1.selected_knot}} in the full series. It is a fitted summary within the registered candidate set, not an independently established physical transition date. The earlier-endpoint search considers a shorter candidate set according to the registered segment-length rule. Thus, the endpoint sensitivity changes both the observations and that candidate set. All covariance-lag sensitivities, the conventional t intervals, exceedance counts and simulation intervals are in the complete result record.

The full-endpoint search gives p-values of {{R1.display.rho_sensitivity.rho_0p2}}, {{R1.display.rho_sensitivity.rho_0p4}} and {{R1.display.rho_sensitivity.rho_0p6}} under the successively stronger registered lag coefficients. This variation illustrates sensitivity of the conditional test to the noise assumption. It does not show that the largest coefficient is the best description of these data.

For the successively larger registered slope changes, the search has simulated conditional power of {{R1.display.power_percent.delta_0p1}}%, {{R1.display.power_percent.delta_0p2}}% and {{R1.display.power_percent.delta_0p3}}%. Their Monte Carlo standard errors and Wilson intervals are supplied in the complete result record. Limited power at the smaller alternatives prevents an absence-of-acceleration conclusion from a non-significant result.

Figure 1 shows the annual series with the fixed-knot fits at all endpoints, and their HAC intervals for the rate increase. The earlier endpoint has a smaller estimated increase and an interval that includes zero.

![Figure 1. Annual temperature anomalies and endpoint sensitivity of the fixed-knot rate increase.](results/endpoint-sensitivity.png)

## Limitations

The analysis concerns one central temperature series and does not propagate observational uncertainty, compare every major temperature dataset, adjust for ENSO, volcanism or solar variability, test an ARMA noise model, or model climate forcing. The fitted AR(1) null is a plug-in approximation. Its parameters are estimated from detrended residuals that may retain a changing trend, so their uncertainty and potential misspecification are not integrated into the simulated p-values. HAC intervals are asymptotic and rely on the selected bandwidth; the record is short for strong claims about covariance estimation. The likelihood and hypothesis tests do not determine the mechanism of any acceleration, its persistence, or future threshold-crossing dates.

The literature informed the fixed knot, the start year and the question. Registration therefore does not make the fixed-knot significance test immune to the broader selection that occurred before this study. Annual aggregation can hide monthly dependence. The chosen endpoint truncations change the length of the post-knot segment; the breakpoint-search candidate set also changes under the registered minimum-length rule. The bootstrap searches test nonlinearity against a straight-line null under specific noise assumptions, rather than every definition of acceleration. A p-value is neither the probability the null is true nor a measure of climate consequence. This audit does not contradict studies using different, adjusted inputs, and makes no first-discovery claim.

## Provenance

The GPT-6 family searched primary literature, designed and registered the analysis, generated the code, ran author cross-checks, interpreted results, wrote the paper and screened the bundle for hazards. The NASA public scientific data are reused with attribution. The code, simulations, derived tables and figure are generated for this audit; no individual-level observations or personal data are used. Python, NumPy, SciPy, statsmodels and Matplotlib supply the numerical and plotting tools. Package versions and the base image digest are pinned. No person wrote the paper or performed calculations. Independent ledger reproduction and review are pending at submission.
