There is no single exact binomial sample size that applies to every study. The required number of observations depends on the null proportion, the alternative proportion, the chosen Type I error rate, the desired Type II error rate, the number of planned looks, and whether the design uses a one-sided or two-sided exact test. The defensible procedure is to define those quantities first and then search for the smallest sample size whose exact test has actual error rates no greater than the stated targets. For a one-sided test of H0: p = p0 against H1: p > p0, the exact p-value is the upper-tail probability P_{p0}(X >= x), where X follows a binomial distribution. For a two-sided exact test, the definition must be selected carefully because equal-tail tests, central tests, and probability-ordering tests can produce different sample sizes. As of 30 September 2026, software can calculate these quantities quickly, but software does not remove the need to make scientifically meaningful assumptions.
What Does “Exact Binomial Sample Size” Actually Mean?
Also worth reading: How Do You Run Book Cover Conversion Testing Without Guessing? · How Should an Exact Binomial Power Design Be Planned for a Small Phase II Study? · How Do You Calculate the Power of a Statistical Cover Test?
An exact binomial calculation uses the discrete probability distribution
P(X = x) = C(n, x) p^x (1-p)^(n-x),
rather than relying on a large-sample normal approximation. If x is the number of observed successes among n independent trials, an upper-tail one-sided exact p-value under a null proportion p0 is P(X >= x | p0). A sample-size calculation answers a different question: for a specified n and decision rule, how often will the test make a Type I error when p = p0, and how often will it fail to reject a false null when p equals a selected alternative such as p1? Those operating characteristics are the important result, not merely the p-value obtained from one hypothetical data set.
“Exact” therefore describes how probabilities and test results are computed; it does not mean that the assumptions are automatically correct or that the resulting answer is exact in an absolute sense. The model still assumes independent Bernoulli trials, a stable success probability, and a sufficiently stable measurement process. If observations come from repeated visits by the same participant, multiple farms on the same research station, duplicate laboratory assays, or a changing population proportion, treating every observation as independent may make both the test and the sample size misleading. Exact arithmetic avoids a normal approximation, but it cannot repair dependence, selection bias, or an unrealistic target difference.
The Inputs That Determine the Sample Size
The first required input is p0, the proportion used in the null hypothesis. A common null value is 0.50, but it should reflect the benchmark relevant to the decision. For example, a publisher may test whether a new headline has a click probability above 10%, an agricultural researcher may test whether seed germination exceeds 80%, or a quality team may test whether a defect rate is below 2%. The second input is p1, the smallest alternative proportion worth detecting. The difference |p1 - p0| is often more informative than either number by itself. A study designed to move a rate from 10% to 14% addresses a much larger contrast than one designed to move it from 10% to 11%, even though the first baseline is identical.
The third input is alpha, the permitted probability of rejecting a true null hypothesis, conventionally written as 0.05. A one-sided test can use the full 0.05 budget because it tests one specified direction. A two-sided test generally uses 0.025 in each tail, although exact discrete testing may behave differently from the familiar equal-tail construction. The fourth input is 1 - beta, the desired probability of rejecting the null when p = p1; power of 80% corresponds to beta = 0.20 and 90% corresponds to beta = 0.10. Fifth, the researcher must state whether the test is evaluated after one analysis or several interim analyses. Repeated testing inflates false-positive probability unless an exact spending function, alpha adjustment, or another control method is specified.
| Design choice | Lower numerical burden | Greater numerical burden | Main caution |
|---|---|---|---|
| One-sided versus two-sided test | One-sided test | Two-sided test | A one-sided test is defensible only when the direction was chosen before seeing data |
| 80% versus 90% power | 80% power | 90% power | A few extra percentage points of power can materially raise n |
| Large versus small p1 - p0 | Larger contrast | Smaller contrast | Small effects may not be practically or operationally important |
| One look versus repeated looks | One final analysis | Interim analyses | Repeated looks require multiplicity control |
| Exact discrete test versus approximation | Often larger n in some settings | Sometimes smaller n | Compare achieved error rates, not only the software’s method label |
Begin by writing a precise hypothesis. For a one-sided quality test, the null might be H0: p = 0.02 and the alternative H1: p > 0.02. Next, choose a scientifically meaningful p1, such as 0.05, and confirm that detecting a rise from 2% to 5% is worth the sample. Set alpha to 0.05 and power to 80% or 90%. Because the test is discrete, define the rejection rule for each possible n. Under H0, find the smallest observed count x that makes the upper-tail exact p-value at most 0.05. A test based on “at least x successes” then has an actual Type I error no greater than 0.05, assuming the null value and independence assumptions are correct.
After defining that rejection rule, calculate power at p1. The rejection probability is P_{p1}(X >= x), using the same threshold x. If that value is below the desired 80%, increase n and repeat the process. If the first feasible n produces power well above 80%, retain the smallest n that meets the criterion and check how conservative the discrete test is. Integer counts create thresholds, so power changes in steps rather than smoothly. For example, adding one trial may not change the minimum qualifying count and may leave achieved power unchanged. A software package can automate this search, but recording the final threshold and achieved error rates is more transparent than reporting only “n = 248.”
A useful spreadsheet or script should iterate over candidate values of n, calculate each exact tail probability, and identify qualifying rules. The same script can evaluate several p1 values and both one- and two-sided constructions. If an interim analysis is planned, the algorithm must simulate or directly calculate the entire decision path, including stopping boundaries; searching for a single-final-analysis sample size is then insufficient. The final protocol should name the software or routine, the test definition, the tails used, alpha, power, p0, p1, and the number of analyses.
A Concrete Example With Actual Numbers
Suppose a diagnostic process is being evaluated with a one-sided exact binomial test. The null proportion is p0 = 0.10, the alternative is p1 = 0.15, alpha is 0.05, and desired power is 80%. Because p0 is exactly 0.10 and n is a multiple of 10, the rejection region P(X >= x | p = 0.10) <= 0.05 can be defined cleanly. Searching over candidate sample sizes generally places the requirement in the low hundreds rather than the few dozen often seen when comparing 10% with 30% or 50%.
A useful sanity check comes from a normal-approximation starting formula for a one-sided comparison of two independent proportions:
n = [z_(1-alpha) sqrt(2 p-bar (1-p-bar)) + z_(1-beta) sqrt(p0(1-p0) + p1(1-p1))]^2 / (p1-p0)^2,
where p-bar = (p0 + p1)/2. For p0 = 0.10, p1 = 0.15, alpha = 0.05, and 80% power, z values of approximately 1.6449 and 0.8416 produce a starting estimate of about 325 observations. The exact search should then be used to establish the final value and threshold, because the approximation does not represent the discreteness or conservatism of the exact test.
The interpretation matters more than the arithmetic. Roughly 325 independent trials are not “325 random facts” that guarantee the observed proportion will be 15%. They provide a design calibrated to attain approximately 80% probability of rejecting the null if the true rate is 15%, subject to the specified rule. If the true rate is 12%, that same study has much lower power because 12% is halfway between 10% and 15%. If the process changes during data collection, if outcomes are correlated, or if a different exact two-sided definition is chosen, the planned operating characteristics may no longer hold.
Exact, Normal, Score, and Bayesian Alternatives
Normal-approximation methods are fast and familiar, but they become less dependable when expected successes or failures are small, when p is near 0 or 1, or when n is not large relative to the hypothesized rates. The standard Wald interval for a proportion, p-hat plus or minus 1.96 sqrt[p-hat(1-p-hat)/n], is especially fragile when np-hat or n(1-p-hat) is below about 5 to 10. Wilson score intervals usually perform better than Wald intervals over a wider range, but an interval method and a hypothesis-test method are not interchangeable merely because both concern binomial proportions.
Exact confidence intervals also require a definition. Clopper–Pearson intervals are produced by inverting exact tests and guarantee at least the stated coverage under repeated sampling, but they can be wider than methods that use a normal approximation. Jeffreys intervals, Bayesian beta-binomial intervals, and other Bayesian credible intervals answer related but distinct questions. A Bayesian design can use a Beta prior and a posterior decision rule; the prior predictive distribution is not the same as a frequentist 95% confidence interval. The exact term alone is therefore too vague for a purchasing specification or protocol.
| Method | How it handles the binomial distribution | Best use | Common weakness |
|---|---|---|---|
| Exact binomial test | Uses discrete tail probabilities | Small samples, boundary proportions, protocol-faithful error control | Conservative and sensitive to the chosen two-sided definition |
| Normal/Wald test | Uses a mean and variance approximation | Large samples away from 0 and 1 | Can produce poor calibration in sparse or extreme cases |
| Score or Wilson procedure | Uses an estimated standard error under the null or improved interval algebra | Common two-proportion and confidence-interval work | Still an approximation in most implementations |
| Bayesian beta-binomial model | Updates a prior on p using observed successes and failures | Decisions with an explicit prior and decision loss | Results depend on prior and loss assumptions |
| Monte Carlo or bootstrap method | Estimates distributions by simulation or resampling | Complex designs, dependence, or custom decision rules | Requires a correct generative model and enough simulations |
The most frequent error is choosing p1 because it sounds favorable rather than because it represents the smallest worthwhile change. Another is reporting a sample size derived from a 50% baseline without stating the target effect. People also substitute the sample-size requirement for testing one proportion against 0.50 with the requirement for comparing two independent proportions, even though the variance and research question differ. Treating p0 as the planning alternative is another error: under H0, the expected number of events is np0, but the probability of detecting the effect is primarily governed by p1.
Discrete testing adds a further source of confusion. There may be no integer n for which the actual Type I error is exactly 0.05 and power exactly 80% at the same time. Investigators should report the chosen rejection threshold, achieved Type I error, achieved power, and conservatism. They should not silently convert an exact two-sided p-value into a doubled upper-tail p-value, because multiplying a discrete tail probability by two can make the nominal level unattainable. In very small samples, several outcomes can share the same p-value, and a test can have Type I error far below the nominal alpha.
Power calculations also fail when observations are not independent. Repeated measurements within patients, several fields from one farm, pages from one website account, or technical replicates from one specimen often reduce the effective information. A correction can sometimes be expressed through a design effect or intraclass correlation, but the better remedy may be to redesign the sampling. A nominally large n of 500 correlated observations can contain less independent information than n = 300 genuinely independent units. Missing observations, treatment nonadherence, and measurement error should be built into the operational target where possible, rather than handled by simply inflating n by an unsupported percentage.
When to Act, Budget, or Seek Expert Help
No paid calculator is necessary for a modest one-sample problem. R, Python, Stata, SAS, SciPy, and several validated online calculators can evaluate binomial distributions and search candidate sample sizes at no direct software cost. The relevant expense is not clicking a button; it is obtaining an adequate number of independent observations and protecting the quality of data collection. If a study has two treatment arms, a cluster-randomized structure, time-to-event outcomes, nonadherence, unequal allocation, subgroup analyses, or multiple interim looks, the basic one-proportion formula is no longer sufficient by itself.
A statistical consultant or trial methodologist is worth consulting when the null proportion is near a regulatory boundary, expected event counts are low, the exact two-sided definition matters, or false-positive control must be demonstrated to an auditor, ethics committee, journal, or regulator. In 2026 dollars, routine independent review may range from roughly US$500 for a narrow calculation check to several thousand dollars for a full protocol, simulation study, or reproducible statistical analysis plan. Complex oncology, agricultural field, platform, or adaptive designs can cost more, but the price should be tied to scope and deliverables rather than sold as an automatic guarantee of exactness.
AI can help generate code, compare assumptions, and draft reproducible documentation, but it should not choose the clinically or operationally meaningful alternative without domain review. An AI-generated answer can also misstate tail conventions, use a normal approximation when an exact test was requested, or overlook clustering. For storywriter.pro’s AI Publishing Consultant context, the sensible service is a transparent workflow in which a consultant records assumptions, runs reproducible calculations, compares methods, and explains limitations; it is not a promise that software produces one universal “exact” number.
The Defensible Reporting Standard
A publication-ready sample-size statement should identify the distribution model, independence assumption, null and alternative proportions, one-sided or two-sided test, exact p-value definition, alpha, desired power, planned number of analyses, and treatment of ties or discrete outcomes. It should also state the rejection threshold, actual Type I error, achieved power, and whether the reported n counts people, trials, fields, specimens, or another independent unit. A strong example is: “A one-sided upper-tail exact binomial test of H0: p = 0.10 against H1: p > 0.10 was evaluated at alpha = 0.05 with 80% power at p = 0.15. The smallest feasible sample size and its operating characteristics were selected by exact enumeration.”
If normal, Bayesian, cluster, or simulation methods are used, they should be labeled accurately and justified rather than placed under the word exact. Researchers should preserve the script, package versions, random seed where simulation is used, and a dated output file. This makes the result auditable when assumptions change and helps another analyst reproduce the number. It also makes clear that a design is conditional: if the true p differs from both p0 and p1, performance will differ, just as a weather forecast’s calibration at one temperature does not guarantee the same accuracy at every temperature.
The direct answer is therefore methodological rather than numerical. To calculate an exact binomial sample size, define H0 and H1, select p1, set alpha and power, choose the exact test construction, enumerate feasible integer sample sizes, and retain the smallest design that meets both error criteria. There is no honest way to supply a definitive n without those inputs. Once they are supplied, the calculation is often inexpensive and can be verified in minutes, but collecting the required independent observations may be the expensive and decisive part of the project.