# How often a 95% interval for a binomial proportion covers it: exact coverage for every sample size from 5 to 100

## Summary

A nominal $95\%$ confidence interval for a binomial proportion should contain the true proportion $p$ in about $95\%$ of samples. Computed exactly, with nothing simulated, for every sample size $5 \le n \le 100$ and every $p$ on a grid of step $0.001$, the textbook Wald interval covers $p$ less than $0.93$ of the time at a fraction {{R1.share_below_low.wald}} of the {{R1.pairs}} $(n, p)$ pairs, and averages {{R1.mean_coverage.wald}}. The Wilson score interval falls below $0.93$ at a fraction {{R1.share_below_low.wilson}} of them, and the Agresti-Coull interval at {{R1.share_below_low.agresti_coull}}. The Clopper-Pearson interval never covers less than $0.95$, but averages {{R1.mean_coverage.clopper_pearson}}, so it is wider than it needs to be. Adding one observation can lower the Wald interval's coverage by more than a tenth. These values reproduce, on a fine grid and with every value declared, the behavior Brown, Cai and DasGupta (2001) described.

## Claims

- **C1.** The Wald interval's coverage is below $0.93$ at a fraction {{R1.share_below_low.wald}} of the pairs, and below the nominal $0.95$ at {{R1.share_below_nominal.wald}}.
- **C2.** The Wilson score interval's coverage is below $0.93$ at a fraction {{R1.share_below_low.wilson}} of the pairs and the Agresti-Coull interval's at {{R1.share_below_low.agresti_coull}}. Their mean coverages are {{R1.mean_coverage.wilson}} and {{R1.mean_coverage.agresti_coull}}, against the Wald interval's {{R1.mean_coverage.wald}}.
- **C3.** The Clopper-Pearson interval's coverage is at least $0.95$ at every pair, at lowest {{R1.min_coverage.clopper_pearson}}, and averages {{R1.mean_coverage.clopper_pearson}}.
- **C4.** At a fixed $p$, the Wald interval's coverage can fall sharply when one observation is added. At $p = 0.2$ its largest fall is from {{R1.wald_at_fixed_p.p_0_2.largest_drop.from}} at n = {{R1.wald_at_fixed_p.p_0_2.largest_drop.from_n}} to {{R1.wald_at_fixed_p.p_0_2.largest_drop.to}} one observation later, and at $p = 0.05$ from {{R1.wald_at_fixed_p.p_0_05.largest_drop.from}} at n = {{R1.wald_at_fixed_p.p_0_05.largest_drop.from_n}} to {{R1.wald_at_fixed_p.p_0_05.largest_drop.to}}.
- **C5.** Even at $n = 100$, the Wald interval covers $p = 0.01$ with probability {{R1.wald_at_fixed_p.p_0_01.at_n_max}} and $p = 0.05$ with probability {{R1.wald_at_fixed_p.p_0_05.at_n_max}}.

## Methods

For a sample size $n$, a true proportion $p$, and $X \sim \mathrm{Binomial}(n, p)$, an interval method's coverage is

$$C(n, p) = \sum_{k=0}^{n} \mathbf{1}\left[L(k) \le p \le U(k)\right] \binom{n}{k} p^k (1-p)^{n-k},$$

where $[L(k), U(k)]$ is the interval the method gives when it sees $k$ successes. [The code](code/coverage.py) evaluates this sum for each method, for every $n$ from 5 to 100 and every $p = i/1000$ with $i$ from 1 to 999, 95,904 pairs in all. Intervals are closed and are not clipped to $[0, 1]$, which changes no coverage since $0 < p < 1$. With $\hat p = k/n$ and $z = 1.959963984540054$, the $0.975$ quantile of the standard normal distribution:

- **Wald:** $\hat p \pm z \sqrt{\hat p (1 - \hat p) / n}$.
- **Wilson score** (Wilson, 1927): $\dfrac{\hat p + z^2/(2n)}{1 + z^2/n} \pm \dfrac{z}{1 + z^2/n} \sqrt{\dfrac{\hat p (1 - \hat p)}{n} + \dfrac{z^2}{4n^2}}$.
- **Agresti-Coull** (Agresti and Coull, 1998): with $\tilde n = n + z^2$ and $\tilde p = (k + z^2/2)/\tilde n$, the interval $\tilde p \pm z \sqrt{\tilde p (1 - \tilde p)/\tilde n}$.
- **Clopper-Pearson** (Clopper and Pearson, 1934): $L(k)$ solves $P(X \ge k \mid p) = 0.025$ for $k \ge 1$, with $L(0) = 0$, and $U(k)$ solves $P(X \le k \mid p) = 0.025$ for $k \le n - 1$, with $U(n) = 1$. Each is found by 80 bisections of $[0, 1]$, past the limit of double precision.

The binomial probabilities come from the ratio of consecutive terms, starting from $P(X = 0) = (1-p)^n$ when $p \le 1/2$ and from $P(X = n) = p^n$ otherwise, each power taken by repeated multiplication. The code uses only addition, subtraction, multiplication, division, and square roots, which IEEE 754 rounds identically everywhere, and sums floats with `math.fsum`, which is correctly rounded, so its results depend neither on the platform's math library nor on the Python version (Python's built-in `sum` compensates since 3.12). It needs only Python's standard library and takes about 20 seconds; `code/run` runs it and writes `results/R1.json`, with every value rounded to six decimals.

Coverage below $0.93$, two points under nominal, counts here as poor; the threshold is a choice, and `results/R1.json` also gives each method's share of pairs below $0.95$, its lowest coverage, and its mean. To follow the Wald interval as $n$ grows, it also records, at $p$ of $0.01$, $0.05$, $0.2$, and $0.5$, how often adding one observation lowers its coverage, the largest such fall, and the largest $n$ at which its coverage is still below $0.93$.

## Results

Over all {{R1.pairs}} pairs:

| Interval | Share below $0.93$ | Share below $0.95$ | Lowest | Mean |
| --- | --- | --- | --- | --- |
| Wald | {{R1.share_below_low.wald}} | {{R1.share_below_nominal.wald}} | {{R1.min_coverage.wald}} | {{R1.mean_coverage.wald}} |
| Wilson score | {{R1.share_below_low.wilson}} | {{R1.share_below_nominal.wilson}} | {{R1.min_coverage.wilson}} | {{R1.mean_coverage.wilson}} |
| Agresti-Coull | {{R1.share_below_low.agresti_coull}} | {{R1.share_below_nominal.agresti_coull}} | {{R1.min_coverage.agresti_coull}} | {{R1.mean_coverage.agresti_coull}} |
| Clopper-Pearson | {{R1.share_below_low.clopper_pearson}} | {{R1.share_below_nominal.clopper_pearson}} | {{R1.min_coverage.clopper_pearson}} | {{R1.mean_coverage.clopper_pearson}} |

The Wald interval's lowest coverage, {{R1.min_coverage.wald}}, comes at the extremes of the grid, where a sample of all failures or all successes gives an interval of width zero that can't contain $p$. Its trouble isn't confined to small samples or extreme $p$: its mean coverage over the grid rises with $n$, from {{R1.by_n.10.wald.mean}} at $n = 10$ to {{R1.by_n.100.wald.mean}} at $n = 100$, still short of nominal. At $p = 0.5$, where it is often assumed to work, adding one observation lowers its coverage at {{R1.wald_at_fixed_p.p_0_5.drops}} of the {{R1.wald_at_fixed_p.p_0_5.steps}} steps from $n = 5$ to $n = 100$, and its coverage is still below $0.93$ at n = {{R1.wald_at_fixed_p.p_0_5.last_n_below_low}}.

Mean coverage at four sample sizes:

| Interval | $n = 10$ | $n = 20$ | $n = 50$ | $n = 100$ |
| --- | --- | --- | --- | --- |
| Wald | {{R1.by_n.10.wald.mean}} | {{R1.by_n.20.wald.mean}} | {{R1.by_n.50.wald.mean}} | {{R1.by_n.100.wald.mean}} |
| Wilson score | {{R1.by_n.10.wilson.mean}} | {{R1.by_n.20.wilson.mean}} | {{R1.by_n.50.wilson.mean}} | {{R1.by_n.100.wilson.mean}} |
| Agresti-Coull | {{R1.by_n.10.agresti_coull.mean}} | {{R1.by_n.20.agresti_coull.mean}} | {{R1.by_n.50.agresti_coull.mean}} | {{R1.by_n.100.agresti_coull.mean}} |
| Clopper-Pearson | {{R1.by_n.10.clopper_pearson.mean}} | {{R1.by_n.20.clopper_pearson.mean}} | {{R1.by_n.50.clopper_pearson.mean}} | {{R1.by_n.100.clopper_pearson.mean}} |

The Wilson interval's mean coverage stays close to nominal at every $n$, but its lowest coverage, {{R1.min_coverage.wilson}}, shows it can dip well below at particular $p$, which is what the share below $0.93$ counts. The Agresti-Coull interval dips less, at the cost of slightly wider intervals, and the Clopper-Pearson interval never dips below nominal, at the cost of intervals wider still. These are the trade-offs that led Brown, Cai and DasGupta to recommend the Wilson interval, or the Jeffreys interval, for small $n$, and the Agresti-Coull interval for larger $n$.

## Limitations

Coverage is computed on a grid of $p$, not at every $p$: coverage jumps where $p$ crosses an interval's endpoint, so the lowest coverage between grid points can be lower than the lowest on the grid. The grid stops $0.001$ short of 0 and 1, so the near-zero coverage of the Wald interval at extreme $p$ appears only at the edges. Means weight every grid point equally, which corresponds to a uniform prior on $p$ rather than to any particular application. The study covers only two-sided intervals at the $95\%$ level, and $n$ only up to 100. It doesn't compute interval widths, so the claim that the conservative intervals are wider rests on their known construction rather than on a measurement here. The Jeffreys interval, which Brown, Cai and DasGupta also recommend, is left out, because its endpoints need the incomplete beta function, which the standard library lacks.

This is a computation of known quantities; its value is in making exact, checkable values available for every $n$ up to 100, not in a new finding.

## Provenance

Written and run by Claude, a model from Anthropic. The code was written for this study, and the results are the code's output, unedited. No data were collected. See `provenance.json` and `references.json`.
