Sample Size for Independent Cohort Studies
Menu location: Analysis_Sample Size_Independent Cohort.
This function gives the minimum number of case subjects required to detect a true relative risk or experimental event rate with power POWER and two sided type I error probability ALPHA. This sample size is also given as a continuity-corrected value intended for use with corrected chi-square and Fisher's exact tests (Casagrande et al. 1978; Meinert 1986; Fleiss, 1981; Dupont, 1990).
Information required
- POWER: probability of detecting a real effect.
- ALPHA: probability of detecting a false effect (two sided: double this if you need one sided).
- P0: probability of event in controls.
- *: input either P1 or RR, where RR=P1/P0.
- P1: probability of event in experimental subjects.
- RR: relative risk of events between experimental subjects and controls.
- M: number of control subjects per experimental subject.
Practical issues
- Usual values for POWER are 80%, 85% and 90%; try several in order to explore/scope.
- 5% is the usual choice for ALPHA.
- P0 can be estimated as the population prevalence of the event under investigation.
- If possible, choose a range of relative risks that you want to have the statistical power to detect.
Technical validation
The estimated sample size n is calculated as:
- where α = alpha, β = 1 - power, nc is the continuity corrected sample size and zp is the standard normal deviate for probability p. n is rounded up to the closest integer.
Example
Suppose you plan a cohort study to find out whether an exposure doubles the risk of a disease. About 10% of unexposed people develop the disease over the period of follow-up, so a relative risk of 2 means that 20% of exposed people would. You will recruit one unexposed subject for each exposed subject, and you want 80% power to detect this relative risk at the 5% two sided level. The figures are invented for this illustration.
To run this in StatsDirect select Independent Cohort from the Sample Size section of the Analysis menu. Enter 0.1 as the probability of the event in the control group, choose the second option (it asks for the relative risk of events between subjects and controls) and enter 2, enter 1 as the number of controls per experimental subject, then 80% as the power and 5% as alpha.
For this example:
Sample size for independent cohort study
Probability of event in control group = 0.1
Probability of event in experimental group = 0.2
Controls per case subject = 1
Alpha = 0.05
Power = 0.8
For uncorrected chi-square test:
N = 199 case subjects and 199 controls
For corrected chi-square and Fisher's exact tests:
N = 219 case subjects and 219 controls
If the groups are to be compared with an uncorrected chi-square test, 199 exposed and 199 unexposed subjects are needed; for a continuity corrected chi-square test or Fisher's exact test, 219 of each. These are the numbers to be followed up, so recruit more to allow for losses.
R code
This R code reproduces the example above. It needs no packages and was checked with R 4.6.1. Paste it into R, or save it as a script and run it.
# Sample size for an independent cohort study: the StatsDirect help example (invented
# figures: 10% of unexposed subjects develop the disease and a relative risk of 2 is to
# be detected, one control per exposed subject, 80% power, 5% two sided alpha) in R
p0 <- 0.1 # probability of the event in the controls
rr <- 2 # the relative risk to detect
p1 <- p0 * rr # probability of the event in exposed subjects
m <- 1 # controls per experimental subject
power <- 0.8
alpha <- 0.05
# power.prop.test with its defaults (two sided, strict = FALSE) uses the same normal
# approximation as the report's uncorrected sample size (the variance pooled under the
# null hypothesis, separate variances under the alternative) but for two groups of
# equal size, so it applies when m = 1; it solves for n numerically and prints the
# unrounded n per group
print(power.prop.test(p1 = p0, p2 = p1, power = power, sig.level = alpha))
# The report's figures, from the formula in the topic, for any number m of controls
# per experimental subject: n is rounded up to a whole number, and the number of
# controls is m times that, rounded down
z_alpha <- qnorm(1 - alpha / 2)
z_beta <- qnorm(power)
pbar <- (p1 + m * p0) / (m + 1)
n <- (z_alpha * sqrt((1 + 1 / m) * pbar * (1 - pbar)) +
z_beta * sqrt(p0 * (1 - p0) / m + p1 * (1 - p1)))^2 / (p0 - p1)^2
cat("Probability of event in control group =", p0, "\n")
cat("Probability of event in experimental group =", p1, "\n")
cat("Controls per case subject =", m, "\n")
cat("Alpha =", alpha, "\n")
cat("Power =", power, "\n")
cat("For uncorrected chi-square test:\n")
cat("N =", ceiling(n), "case subjects and", floor(m * ceiling(n)), "controls\n")
# The continuity-corrected size for the corrected chi-square and Fisher's exact tests
# (Casagrande et al. 1978; Fleiss 1981), calculated from the unrounded n
nc <- n / 4 * (1 + sqrt(1 + 2 * (m + 1) / (n * m * abs(p0 - p1))))^2
cat("For corrected chi-square and Fisher's exact tests:\n")
cat("N =", ceiling(nc), "case subjects and", floor(m * ceiling(nc)), "controls\n")