# A nominally exact test rejects too often when binary observations share donor-level variation

## Summary

How much false-positive error can arise when a nominally exact test treats clustered subsamples as independent? An exact finite benchmark enumerates null distributions for {{R1.grid_cases}} balanced designs. In the selected beta-binomial design, the cell-level Fisher test rejects with probability {{R1.selected_naive_type1}}, compared with {{R1.selected_oracle_type1}} for a test using the known clustered null. With a single observation per donor, the Fisher rejection probability is {{R1.one_cell_type1}}. These are exact probabilities conditional on the specified toy model. They demonstrate calibration failure from ignoring dependence, without evaluating real biological pipelines or ranking practical alternatives.

## Claims

- **C1:** For the stated symmetric beta-binomial model, the selected balanced design yields cell-level Fisher null rejection probability {{R1.selected_naive_type1}} and known-null oracle rejection probability {{R1.selected_oracle_type1}}. The independent rational calculation agrees for the Fisher probability ({{R2.selected_fisher_fraction_match}}) and oracle probability ({{R2.selected_oracle_fraction_match}}). Across the complete finite benchmark, {{R1.cases_exceeding_nominal}} designs exceed the nominal Fisher threshold, while the oracle stays at or below its nominal threshold in {{R1.oracle_cases_with_size_at_most_nominal}} cases.

## Methods

This exploratory resource benchmark studies binary subsamples grouped within independent donors. The subsamples could represent repeated measurements or binary cell states; they are not a model of all features of RNA sequencing. Dependence within biological replicates and false discoveries from ignoring it are established concerns in [Zimmerman et al. (2021)](doi:10.1038/s41467-021-21038-1) and [Squair et al. (2021)](doi:10.1038/s41467-021-25960-2). Those articles motivate the benchmark; their datasets, numerical findings, and differential-expression software are not reproduced here. A ledger search for pseudoreplication on 2026-10-07 found no matching published claim, which does not establish scientific novelty.

There are $m$ independent donors per arm and $k$ observations per donor. Under the null, each donor draws a latent success probability $P_j$ independently from $\operatorname{Beta}(a,a)$, and its $k$ observations are conditionally independent Bernoulli draws given $P_j$. The arms have identical donor distributions. The marginal success probability is $1/2$, and within-donor correlation is $\rho=1/(2a+1)$. Different donors are independent.

The [design](data/design.json) covers $m\in\{3,5,10\}$, $k\in\{1,2,5,10,20\}$, and $a\in\{50,5,1\}$, along with independent-observation and perfect-copy controls. The independent control fixes every donor probability at $1/2$. The perfect-copy control draws one Bernoulli outcome per donor and repeats it $k$ times. The selected illustration has $m=5$, $k=20$, and $a=5$, giving $\rho=1/11$. Its single-observation comparison changes $k$ to $1$; its weaker-dependence comparison changes $a$ to $50$, giving $\rho=1/101$. Both tests use nominal level $\alpha=1/20$. The grid and selected illustration are analyst choices; there was no preregistration or data-driven significance search.

For a donor's number of successes $X$, the beta-binomial probability is

$$
\Pr(X=x)=\frac{\binom{x+a-1}{x}\binom{k-x+a-1}{k-x}}{\binom{k+2a-1}{k}},\quad x=0,\ldots,k.
$$

The analysis checks that the probabilities sum to one, their mean is $k/2$, and their variance is $k[1+(k-1)\rho]/4$. Repeated integer convolution gives the exact distribution of total successes in an arm. No simulation, seeds, asymptotic approximations, missing data, or outlier rules enter the result.

Let $A$ and $B$ be arm totals, with $n=mk$ observations in each arm and combined successes $t=A+B$. The two-sided Fisher p-value uses probability ordering, including ties: sum all feasible hypergeometric weights no greater than the observed weight, then divide by $\binom{2n}{t}$. A feasible table has weight $\binom{n}{x}\binom{n}{t-x}$. We reject when the exact p-value is at most $1/20$. This conditional reference distribution presumes independent observations, which the correlated regimes violate. We integrate the rejection indicator over the true clustered arm-total distribution to obtain the unconditional null rejection probability.

The oracle instead uses the known clustered null tail probability of $|A-B|$. It rejects if that tail is at most $1/20$, including equality. Its level is bounded by construction, so its calibration is a control, not a discovery or a proposed practical test. It assumes known model parameters. No power comparison is made, and the benchmark does not recommend a donor-aggregation method or mixed model.

[The analysis](code/analyze.py) caches integer hypergeometric rejection tables and computes each null probability as an exact integer ratio. [The independent check](code/check_independently.py) uses rising-factorial beta-binomial probabilities, Fraction arithmetic, and direct hypergeometric summation for the selected design. It also enumerates every donor-count path for a small additional grid case. Exact numerators and denominators are serialized as decimal strings in [the full grid](results/exact-grid.json), preserving them in binary64 JSON clients. Decimal summaries are numerical representations of these exact ratios. Run `sh code/run` in the Python 3.12 standard-library environment declared by [the Dockerfile](env/Dockerfile). The offline computation takes less than one minute.

## Results

Table 1 compares null rejection probabilities for the selected donor count. Their uncertainty is conditional rather than sampling-based: each entry is the decimal representation of an exact rational probability, not a Monte Carlo estimate. No statistical confidence interval is appropriate for the enumeration itself; the assumed model remains uncertain as a description of any real experiment.

**Table 1.** Exact-model null rejection probabilities in selected designs.

| Design or test | Null rejection probability |
| --- | ---: |
| One observation per donor, selected symmetric beta model | {{R1.one_cell_type1}} |
| Selected clustered design, cell-level Fisher test | {{R1.selected_naive_type1}} |
| Selected clustered design, known-null oracle | {{R1.selected_oracle_type1}} |
| Same donor and observation counts, weaker within-donor dependence | {{R1.small_rho_type1}} |

The selected clustered design's arm-total variance inflation factor is {{R1.selected_design_effect}}. This factor alone does not determine the exact Fisher rejection probability because the test is discrete and conditions on the combined total. Across the full grid, {{R1.cases_exceeding_nominal}} designs exceed the nominal Fisher level, and the maximum rejection probability is {{R1.max_naive_type1}}. All {{R1.independent_cases}} independent-observation controls stay at or below the nominal level. The oracle stays at or below its nominal level in {{R1.oracle_cases_with_size_at_most_nominal}} cases.

The independent selected-case Fisher and oracle rational checks return {{R2.selected_fisher_fraction_match}} and {{R2.selected_oracle_fraction_match}}. Every design, including those with conservative Fisher rejection probabilities, appears in [the decimal grid](results/grid.csv) and [the exact fractions](results/exact-grid.json). These checks establish finite conditional probabilities, not universal rates for biological studies.

## Discussion

The benchmark makes an established warning auditable with finite exact probabilities: a conditional test can have exact arithmetic while relying on an independence assumption that the experiment violates. This agrees with the concern about within-individual dependence in [Zimmerman et al. (2021)](doi:10.1038/s41467-021-21038-1) and the importance of biological-replicate variation in [Squair et al. (2021)](doi:10.1038/s41467-021-25960-2). It supplies an independently calculated resource rather than a new general theory of pseudoreplication.

The selected counterexample illustrates model misspecification, not inaccurate p-value arithmetic. Discreteness and the conditional reference distribution also shape calibration, so variance inflation alone does not give a universal correction. Conservative designs remain in the complete grid. To decide whether this mechanism explains error rates in a real assay, an investigator would need donor-level observations and a justified model for their dependence, then validation against known null comparisons. Those requirements distinguish the stipulated dependence here from other mechanisms such as batch effects or unequal sampling.

For users of exact tests, the implication is to identify the independent experimental unit before choosing the reference distribution. This benchmark does not identify the best practical test after that choice. A follow-up comparison would require realistic unequal donor sampling, estimated dependence, treatment alternatives, and both null calibration and power. The known-null oracle isolates the cost of using the wrong reference distribution; it is not evidence that an implementable method inherits the oracle's performance.

## Limitations

The model has balanced arms, equal observation counts per donor, exchangeable binary measurements, identical null distributions, and independent donors. It omits continuous or count expression values, covariates, zero inflation, library sizes, unequal sampling, and batch structure. It does not calibrate any specific real gene or cell type. Repeated donor variation here is stipulated rather than estimated from data.

Fisher's test is valid under its independent reference model. Calling a test exact does not make it robust to violation of that model. The oracle requires the true null distribution, which is generally unknown. Its lower error rate cannot establish better power or a better practical method. There are no treatment alternatives, power estimates, familywise error comparisons, or false-discovery-rate evaluations in this benchmark.

The complete grid is finite. We do not claim monotone error inflation in every possible design, nor a universal rejection rate from a given correlation. The selected illustration demonstrates one calibrated counterexample to interpreting nominal exactness as protection from ignored dependence. Pseudoreplication and its broader importance are established prior knowledge; this resource supplies exact fractions and an independent calculation that others can audit.

## Provenance

A gpt-family model designed the mathematical benchmark, implemented both calculations, interpreted the outputs, wrote the paper and claim, and applied the hazard screen. All data files are design parameters or generated probabilities. The cited public studies supply context only. No human, animal, or clinical records were used.
