Chapters 1–5
Formula and concept review
Data summaries, probability,
and discrete models
EN.553.211 · Probability & Statistics for the Life Sciences
FROM OBSERVATIONS TO MODELSDataEventsModels
xˉ, s2\bar{x},\ s^2
P(A)P(A)
E(X),V(X)E(X),V(X)
Name the object → choose the method
01

Data, events, and random variables

Observed data
x1,…,xnx_1,\ldots,x_n
An event
A⊆S\textcolor{#8bded6}{A}\subseteq S
A random variable
X:S⟶R\textcolor{#f4dc88}{X}:S\longrightarrow\mathbb{R}
SS
outcomes
AA
XX
x1x_1
x2x_2
values
Several outcomes can give one value

Meaning

Categorical variables label groups. Quantitative variables measure amounts. An event collects outcomes. A random variable assigns a number to each outcome.

When to use it

“Describe the data” calls for a statistic. “How likely?” calls for an event. “Number of…” often defines a count X.

Population: μ,σ2\mu,\sigma^2 · Sample: xˉ,s2\bar{x},s^2 · Relative frequency = category count / nn
Describing data02

Mean and median

Arithmetic mean
xˉ=1n∑i=1nxiμ=1N∑i=1Nxi\bar{x}=\frac1n\sum_{i=1}^n x_i\qquad \mu=\frac1N\sum_{i=1}^N x_i
Sorted data, odd n; j = (n+1)/2(n+1)/2
x~=xj\textcolor{#8bded6}{\tilde{x}}=x_j
Sorted data, even n; j = n/2n/2
x~=xj+xj+12\textcolor{#8bded6}{\tilde{x}}=\frac{x_j+x_{j+1}}2
xˉ\bar{x}
x~\tilde{x}
even n: average the middle two
Balance point and middle position

Meaning

The mean is the arithmetic balance point. After sorting, xjx_j is the jjth observation and the median x~\tilde{x} divides the data into lower and upper halves.

When to use it

Use the mean for an arithmetic average. Use the median for a resistant center when skew or extreme observations matter.

Skew follows the long tail. Mean–median ordering is a typical pattern, not a universal rule.
Describing data03

Changes to a dataset

Add a new observation y
xˉnew=nxˉ+yn+1\bar{x}_{\mathrm{new}}=\frac{\textcolor{#f4dc88}{n\bar{x}}+\textcolor{#8bded6}{y}}{n+1}
Remove an observed value y, with n>1n>1
xˉnew=nxˉ−yn−1\bar{x}_{\mathrm{new}}=\frac{n\bar{x}-y}{n-1}
Replace an observed x by y
xˉnew=xˉ+y−xn\bar{x}_{\mathrm{new}}=\bar{x}+\frac{y-x}{n}
nxˉn\bar{x}
yy
updated totalupdated count\frac{\text{updated total}}{\text{updated count}}
Remove: subtract, then reduce the countReplace: change the total; keep n fixed
Track the total and the count separately

Meaning

The original total is nxˉn\bar{x}. Update that total and the number of observations, then divide to find the new mean.

When to use it

Use for adding, removing, or replacing data values. To update the median, re-sort and locate the new middle position or positions.

The median depends on order. The original mean and sample size alone do not determine the new median.
Describing data04

Quartiles and the five-number summary

Ordered observations
x1≤x2≤⋯≤xnx_1\le x_2\le\cdots\le x_n
Course quartile positions
n1=n+14n3=3(n+1)4\textcolor{#8bded6}{n_1}=\frac{n+1}{4}\qquad\textcolor{#ee9daf}{n_3}=\frac{3(n+1)}4
Interpolation at position j + d
Q=xj+d(xj+1−xj)Q=x_j+\textcolor{#f4dc88}{d}(x_{j+1}-x_j)
min⁡\min
Q1Q_1
x~\tilde{x}
Q3Q_3
max⁡\max
n1n_1
n3n_3
xjx_j
xj+1x_{j+1}
dd
A fractional position interpolates between neighbors

Meaning

The five-number summary is minimum, Q1Q_1, median x~\tilde{x}, Q3Q_3, maximum. The quartiles mark the lower and upper quarters.

When to use it

For a noninteger position, move fraction dd of the way from observation jj to observation j+1j+1. At an integer position, use that observation.

This is the course convention (n≥3n\ge3). Software can use different quartile definitions.
Describing data05

IQR and outlier boxplots

Two measures of spread
IQR=Q3−Q1range=max⁡−min⁡\textcolor{#f4dc88}{\mathrm{IQR}}=\textcolor{#ee9daf}{Q_3}-\textcolor{#8bded6}{Q_1}\qquad \mathrm{range}=\max-\min
Lower fence
L=Q1−1.5 IQRL=\textcolor{#8bded6}{Q_1}-1.5\,\mathrm{IQR}
Upper fence
U=Q3+1.5 IQRU=\textcolor{#ee9daf}{Q_3}+1.5\,\mathrm{IQR}
Q1Q_1
Q3Q_3
IQR\mathrm{IQR}
LL
UU
whiskers = observed values inside fences
Box, whiskers, fences, and flagged points

Meaning

The box runs from Q1Q_1 to Q3Q_3, with a line at the median. Whiskers end at the most extreme observed values inside the fences.

When to use it

Use a boxplot to compare center, middle-half spread, asymmetry, and possible outliers across quantitative distributions.

Fences are cutoffs. Whiskers are observed values. Flagged observations need investigation.
Describing data06

Population and sample variance

Complete population of N values
σ2=∑i=1N(xi−μ)2N=∑i=1Nxi2N−μ2\sigma^2=\frac{\sum_{i=1}^N(x_i-\mu)^2}{N}=\frac{\sum_{i=1}^Nx_i^2}{N}-\mu^2
Sample of n>1n>1 values
s2=∑i=1n(xi−xˉ)2n−1\textcolor{#f4dc88}{s^2}=\frac{\sum_{i=1}^n\textcolor{#8bded6}{(x_i-\bar{x})^2}}{n-1}
Computational form
s2=∑i=1nxi2−nxˉ 2n−1s^2=\frac{\sum_{i=1}^nx_i^2-n\bar{x}^{\,2}}{n-1}
xˉ\bar{x}
(xi−xˉ)2(x_i-\bar{x})^2
square → add → divide
Squaring makes every deviation contribute positively

Meaning

Variance averages squared deviations from the relevant mean. SD is the square root: σ=σ2\sigma=\sqrt{\sigma^2} and s=s2s=\sqrt{s^2}.

When to use it

Choose the denominator from the task: a complete population, an estimated population variance, or a probability model.

Variance has squared units. SD has original units. An empirical distribution uses 1/n1/n weights.
Describing data07

Shifts, scales, and standardized values

Transform every observation
yi=axi+by_i=\textcolor{#8bded6}{a}x_i+\textcolor{#ee9daf}{b}
Center and spread
yˉ=axˉ+bsy2=a2sx2\bar{y}=a\bar{x}+b\qquad s_y^2=a^2s_x^2
Standardized distance: population / sample
z=x−μσz=x−xˉs\textcolor{#f4dc88}{z}=\frac{x-\mu}{\sigma}\qquad z=\frac{x-\bar{x}}s
original
+b+b
same spread
×a\times a
SD scales by |a|
Shift moves the center; scale changes the spread

Meaning

A shift changes the center. Scaling by aa multiplies SD by ∣a∣|a| and variance by a2a^2. A zz-score measures distance from the mean in SD units.

When to use it

Use these rules for unit conversions, adding a constant, or comparing relative position. Standardization requires positive SD.

For an approximately normal shape: about 68%, 95%, 99.7% lie within 1, 2, 3 SDs.
Describing data08

Covariance

Sample paired observations
sxy=∑i=1n(xi−xˉ)(yi−yˉ)n−1\textcolor{#f4dc88}{s_{xy}}=\frac{\sum_{i=1}^n\textcolor{#8bded6}{(x_i-\bar{x})}\textcolor{#ee9daf}{(y_i-\bar{y})}}{n-1}
Complete paired population
σxy=1N∑i=1N(xi−μx)(yi−μy)\sigma_{xy}=\frac1N\sum_{i=1}^N(x_i-\mu_x)(y_i-\mu_y)
Computational form
sxy=∑i=1nxiyi−nxˉyˉn−1s_{xy}=\frac{\sum_{i=1}^nx_iy_i-n\bar{x}\bar{y}}{n-1}
++−−
xi−xˉx_i-\bar{x}
yi−yˉy_i-\bar{y}
same signs → positive product
Centered coordinates reveal each contribution

Meaning

Covariance averages products of paired deviations. Same-direction deviations contribute positively. Units are xx-units times yy-units.

When to use it

Use it to quantify joint direction, compute correlation, or build the least-squares slope. Keep each observed pair intact.

A scatterplot reveals shape and unusual points that a single number can hide.
Association and prediction09

Correlation

Standardized covariance
r=sxysxsy\textcolor{#f4dc88}{r}=\frac{s_{xy}}{\textcolor{#8bded6}{s_x}\textcolor{#ee9daf}{s_y}}
Population counterpart
ρ=σxyσxσy\rho=\frac{\sigma_{xy}}{\sigma_x\sigma_y}
Both SDs must be positive
−1≤r,ρ≤1∣r∣=1: exact linear association-1\le r,\rho\le1\qquad |r|=1:\ \text{exact linear association}
r<0r<0
r=0r=0
r>0r>0
The middle pattern is nonlinear despite zero correlation

Meaning

Correlation gives the direction and strength of linear association. It is unitless. A larger ∣r∣|r| means a closer linear relationship.

When to use it

Use rr to compare linear association across scales. Inspect the scatterplot for curvature and influential observations first.

With positive SDs: sign⁡(r)=sign⁡(sxy)=\operatorname{sign}(r)=\operatorname{sign}(s_{xy})= sign of the centered cross-product sum.
r=0r=0 does not imply independence. Correlation alone does not establish causation.
Association and prediction10

Least-squares regression

Predict y using x
y^=mx+b\textcolor{#f4dc88}{\hat{y}}=\textcolor{#8bded6}{m}x+b
Slope
m=sxysx2=rsysx\textcolor{#8bded6}{m}=\frac{s_{xy}}{s_x^2}=r\frac{s_y}{s_x}
Intercept (prediction at x=0x=0) and residual
b=yˉ−mxˉei=yi−y^ib=\bar{y}-m\bar{x}\qquad\textcolor{#ee9daf}{e_i}=y_i-\hat{y}_i
xx
y^=mx+b\hat{y}=mx+b
Δx\Delta x
mΔxm\Delta x
eie_i
(xˉ,yˉ)(\bar{x},\bar{y})
Residual = observed minus fitted

Meaning

xx is the explanatory variable and yy is the response. The slope gives predicted yy change per xx unit. A positive residual means observed yy exceeds predicted yy.

When to use it

Use the line for linear prediction with variation in xx. Least squares minimizes ∑iei2\sum_i e_i^2 and the fitted line passes through (xˉ,yˉ)(\bar{x},\bar{y}).

Slope requires sx>0s_x>0. The rr form also needs sy>0s_y>0. Extrapolation needs justification.
Association and prediction11

Events and set operations

Both events occur
A∩B\textcolor{#8bded6}{A}\cap\textcolor{#ee9daf}{B}
At least one occurs
A∪BA\cup B
A occurs without B
A∩BcA\cap B^c
SS
AA
BB
The highlighted region follows the current formula

Meaning

SS is the sample space. An event is a subset of SS. AcA^c contains outcomes outside AA. Disjoint events have A∩B=∅A\cap B=\varnothing.

When to use it

“And” suggests an intersection. Inclusive “or” suggests a union. “Neither” means Ac∩BcA^c\cap B^c. “Exactly one” excludes the overlap.

(A∪B)c=Ac∩Bc(A∩B)c=Ac∪Bc(A\cup B)^c=A^c\cap B^c\qquad(A\cap B)^c=A^c\cup B^c
Events and counting12

Probability rules

Basic constraints
P(S)=10≤P(A)≤1P(S)=1\qquad0\le P(A)\le1
Complement
P(Ac)=1−P(A)P(A^c)=1-P(A)
Union
P(A∪B)=P(A)+P(B)−P(A∩B)P(A\cup B)=\textcolor{#8bded6}{P(A)}+\textcolor{#ee9daf}{P(B)}-\textcolor{#f4dc88}{P(A\cap B)}
SS
− P(A∩B)-\,P(A\cap B)
AA
BB
Subtract the overlap counted twice

Meaning

The union formula subtracts the overlap counted twice. For disjoint events, the overlap is empty and probabilities add.

When to use it

“At least one” is often easiest as one minus “none.” For finite equally likely outcomes only, P(A)=#A/#SP(A)=\#A/\#S.

P(A∩Bc)=P(A)−P(A∩B)P(A\cap B^c)=P(A)-P(A\cap B)
Events and counting13

Counting by the sampling mechanism

Ordered, with replacement
nkeach position has n choicesn^k\qquad\text{each position has }n\text{ choices}
Ordered, without replacement: available choices decrease
(n)k=n!(n−k)!(n)_k=\frac{n!}{(n-k)!}
Unordered, without replacement: remove k!k! reorderings
(nk)=n!k!(n−k)!\binom{n}{k}=\frac{n!}{k!(n-k)!}
nn
nn
nn
ordered positions
nn
n−1n-1
n−2n-2
÷ k!\div\ k!
one group, many orderings
Replacement and order determine the count

Meaning

The product rule multiplies the number of choices at successive stages. Add counts when alternatives are disjoint.

When to use it

“Arrange” or “sequence” makes order matter. “Choose a group” usually makes order irrelevant. Replacement keeps objects available.

0!=1(nk)=(nn−k)0!=1\qquad\binom{n}{k}=\binom{n}{n-k} · Numerator and denominator must count the same objects.
Events and counting14

Conditional probability and multiplication

Restriction to a reference event
P(A∣B)=P(A∩B)P(B)P(A\mid\textcolor{#ee9daf}{B})=\frac{\textcolor{#f4dc88}{P(A\cap B)}}{\textcolor{#ee9daf}{P(B)}}
Multiplication rule
P(A∩B)=P(B)P(A∣B)P(A\cap B)=P(B)P(A\mid B)
Several stages
P(A∩B∩C)=P(A)P(B∣A)P(C∣A∩B)P(A\cap B\cap C)=P(A)P(B\mid A)P(C\mid A\cap B)
SS
P(A∩B)P(B)\frac{P(A\cap B)}{P(B)}
AA
BB
The given event B is the new reference group

Meaning

Conditioning keeps outcomes in BB and rescales their total probability to one. The denominator is the group named after “given.”

When to use it

Use conditional probabilities when earlier outcomes change later chances, or when the question restricts the reference group.

Conditioning events need positive probability. The multiplication rule does not require independence.
Conditioning15

Independence

Definition for two events
P(A∩B)=P(A)P(B)P(A\cap B)=\textcolor{#8bded6}{P(A)}\textcolor{#ee9daf}{P(B)}
Equivalent when P(B)>0P(B)>0
P(A∣B)=P(A)P(A\mid B)=P(A)
Mutual independence: every subcollection
P(Ai1∩⋯∩Aik)=∏j=1kP(Aij)P(A_{i_1}\cap\cdots\cap A_{i_k})=\prod_{j=1}^kP(A_{i_j})
P(A)P(A)
P(B)P(B)
P(A∣B)=P(A)P(A\mid B)=P(A)
same fraction inside B
An independent area model: overlap equals a product

Meaning

Independence means learning one event occurred does not change the probability of the other. Disjointness means both cannot occur.

When to use it

Use a product of marginal probabilities only when independence follows from the model or has been established.

Positive-probability disjoint events are dependent. Pairwise independence alone is insufficient for 3 or more events.
Conditioning16

Law of total probability

A partition of the sample space
⋃i=1mBi=SBi∩Bj=∅ (i≠j)\bigcup_{i=1}^{m}B_i=S\qquad B_i\cap B_j=\varnothing\ (i\ne j)
Probability over all routes
P(A)=∑i=1mP(A∣Bi)P(Bi)\textcolor{#f4dc88}{P(A)}=\sum_{i=1}^{m}\textcolor{#8bded6}{P(A\mid B_i)}\textcolor{#ee9daf}{P(B_i)}
B1B_1
AA
AcA^c
B2B_2
AA
AcA^c
B3B_3
AA
AcA^c
multiply along routes; add across routes
Every group contributes one route to A

Meaning

Each term is the probability of one route into AA: first belong to group BiB_i, then satisfy AA within that group.

When to use it

Use when a population or experiment splits into distinct groups, mechanisms, or starting states with different conditional chances.

Use exhaustive, nonoverlapping groups. Omit zero-probability groups when writing conditionals.
Conditioning17

Bayes’ rule

Reverse the conditioning direction
P(Bj∣A)=P(A∣Bj)P(Bj)P(A)P(B_j\mid A)=\frac{\textcolor{#8bded6}{P(A\mid B_j)}\textcolor{#ee9daf}{P(B_j)}}{\textcolor{#f4dc88}{P(A)}}
Expand the evidence probability
P(Bj∣A)=P(A∣Bj)P(Bj)∑i=1mP(A∣Bi)P(Bi)P(B_j\mid A)=\frac{P(A\mid B_j)P(B_j)}{\sum_{i=1}^{m}P(A\mid B_i)P(B_i)}
B1B_1
AA
AcA^c
B2B_2
AA
AcA^c
B3B_3
AA
AcA^c
retain A routes → renormalize
Which starting group produced the evidence?

Meaning

Prior P(Bj)P(B_j): group probability before evidence. Likelihood P(A∣Bj)P(A\mid B_j): evidence probability within the group. Posterior: group probability after evidence.

When to use it

Use when evidence is observed and the question asks which group or source produced it. Work from the known conditioning direction.

P(A)>0P(A)>0. The denominator includes every route that can produce AA.
Conditioning18

Discrete random variables and PMFs

Probability mass function
f(x)=P(X=x)≥0∑all xf(x)=1\textcolor{#f4dc88}{f(x)}=P(X=x)\ge0\qquad\sum_{\text{all }x}f(x)=1
Event probability: sum over support values in A
P(X∈A)=∑x∈Af(x)P(X\in A)=\sum_{x\in A}f(x)
Cumulative distribution function
F(t)=P(X≤t)=∑x≤tf(x)\textcolor{#8bded6}{F(t)}=P(X\le t)=\sum_{x\le t}f(x)
xx
f(x)f(x)
00
⋯\cdots
∑x∈Af(x)\sum_{x\in A}f(x)
tt
Add only the masses in the requested event

Meaning

The support lists possible values with positive mass. Multiple experimental outcomes can contribute to the same value of XX.

When to use it

Define XX in words, list its support, and translate the question into a set of values. Then sum the corresponding masses.

Capital XX is the random variable. Lowercase xx is a possible value.
Random-variable summaries19

Expectation

Weighted mean
μX=E(X)=E[X]=∑all xxf(x)\textcolor{#f4dc88}{\mu_X}=E(X)=E[X]=\sum_{\text{all }x}\textcolor{#8bded6}{x}f(x)
A function of X
E[g(X)]=∑all xg(x)f(x)E[g(X)]=\sum_{\text{all }x}g(x)f(x)
Second moment
E(X2)=∑all xx2f(x)E(X^2)=\sum_{\text{all }x}x^2f(x)
xx
f(x)f(x)
00
⋯\cdots
μX\mu_X
x⟼g(x)x\longmapsto g(x)
same probability weights
Expectation is the probability-weighted balance point

Meaning

Expectation is a probability-weighted average. E[g(X)]E[g(X)] averages the transformed values using the original probabilities.

When to use it

Use for a long-run average, expected count, or average value after a transformation. The mean need not be a possible outcome.

E[g(X)]E[g(X)] generally differs from g(E(X))g(E(X)). Infinite-support formulas need the relevant moments to exist.
Random-variable summaries20

Variance of a random variable

Definition
σX2=V(X)=Var⁡(X)=E ⁣[(X−μX)2]\textcolor{#f4dc88}{\sigma_X^2}=V(X)=\operatorname{Var}(X)=E\!\left[(X-\mu_X)^2\right]
PMF form
V(X)=∑all x(x−μX)2f(x)V(X)=\sum_{\text{all }x}\textcolor{#8bded6}{(x-\mu_X)^2}f(x)
Computational form
V(X)=E(X2)−[E(X)]2V(X)=E(X^2)-[E(X)]^2
xx
f(x)f(x)
00
⋯\cdots
μX\mu_X
(x−μX)2(x-\mu_X)^2
weight by f(x)
Weight squared distances by probability

Meaning

Variance is the weighted mean squared distance from the mean. It is nonnegative. σX=V(X)\sigma_X=\sqrt{V(X)} returns to the original units.

When to use it

Use the definition to interpret spread. Use the second-moment identity when E(X)E(X) and E(X2)E(X^2) are easier to compute.

Finite second moment assumed. Model probability weights require no n−1n-1 correction.
Random-variable summaries21

Linear combinations: means and variances

Linearity of expectation
E(aX+bY+c)=aE(X)+bE(Y)+cE(aX+bY+c)=aE(X)+bE(Y)+c
Variance with dependence retained
V(aX+bY+c)=a2V(X)+b2V(Y)+2abCov⁡(X,Y)\begin{aligned}V(aX+bY+c)&=a^2V(X)+b^2V(Y)\\&\quad+\textcolor{#ee9daf}{2ab\operatorname{Cov}(X,Y)}\end{aligned}
aE(X)aE(X)
bE(Y)bE(Y)
+means add
a2V(X)a^2V(X)
b2V(Y)b^2V(Y)
+2abCov⁡(X,Y)+2ab\operatorname{Cov}(X,Y)
Dependence contributes the cross term

Meaning

Expectations add without independence. Variance also contains a covariance term. A constant shift contributes no variance.

When to use it

Use for totals, differences, weighted combinations, and changes of units. Retain the covariance term unless it is zero.

For finite second moments: V(aX+c)=a2V(X)V(aX+c)=a^2V(X) and σaX+c=∣a∣σX\sigma_{aX+c}=|a|\sigma_X.
Random-variable summaries22

Covariance, sums, and averages

Model covariance
Cov⁡(X,Y)=E(XY)−E(X)E(Y)\operatorname{Cov}(X,Y)=E(XY)-E(X)E(Y)
Variance of a sum: every pair i<ji<j
V ⁣(∑iXi)=∑iV(Xi)+2∑i<jCov⁡(Xi,Xj)V\!\left(\sum_i X_i\right)=\sum_iV(X_i)+\textcolor{#ee9daf}{2\sum_{i<j}\operatorname{Cov}(X_i,X_j)}
Independent, identically distributed values
E(X‾)=μV(X‾)=σ2nE(\overline{X})=\mu\qquad\textcolor{#f4dc88}{V(\overline{X})}=\frac{\sigma^2}{n}
X1X_1
X2X_2
X3X_3
X4X_4
include each covariance pair once
X1+⋯+XnX_1+\cdots+X_n
X‾=1n∑iXiσX‾=σn\overline{X}=\frac1n\sum_iX_i\quad\sigma_{\overline{X}}=\frac{\sigma}{\sqrt{n}}
For iid values, averaging reduces variance by n

Meaning

Independent variables have E(XY)=E(X)E(Y)E(XY)=E(X)E(Y) and zero covariance. Zero covariance alone does not establish independence.

When to use it

Use the sum formula to track dependence. For an average of nn independent equal-variance measurements, σX‾=σ/n\sigma_{\overline{X}}=\sigma/\sqrt{n}.

E ⁣(∑iaiXi)=∑iaiE(Xi)E\!\left(\sum_i a_iX_i\right)=\sum_i a_iE(X_i) · No independence is needed for this mean formula.
Random-variable summaries23

Bernoulli variables and indicators

One binary outcome
I={1on A0on AcI=\begin{cases}1&\text{on }A\\0&\text{on }A^c\end{cases}
Success and failure probabilities
P(I=1)=pP(I=0)=q=1−pP(I=1)=\textcolor{#8bded6}{p}\qquad P(I=0)=\textcolor{#ee9daf}{q}=1-p
Moments
E(I)=pV(I)=pqE(I)=p\qquad V(I)=pq
00
11
qq
pp
p+q=1p+q=1
one trial, two possible values
The bar heights are schematic: p and q sum to one

Meaning

An indicator records whether an event occurs. “Success” is simply the outcome chosen for counting, and 0≤p≤10\le p\le1.

When to use it

Use Bernoulli for one yes/no outcome. Write a count as X=∑iIiX=\sum_i I_i to compute E(X)=∑iP(Ai)E(X)=\sum_iP(A_i), even with dependence.

A sum of indicators requires extra assumptions to have a binomial distribution.
Discrete count models24

Independent indicators with different probabilities

One success set J; qi=1−piq_i=1-p_i
w(J)=∏i∈Jpi∏1≤i≤ni∉Jqiw(J)=\prod_{i\in J}\textcolor{#8bded6}{p_i}\prod_{\substack{1\le i\le n\\i\notin J}}\textcolor{#ee9daf}{q_i}
Sum over sets J⊆{1,…,n}J\subseteq\{1,\ldots,n\} with ∣J∣=k|J|=k
P(X=k)=∑J⊆{1,…,n}∣J∣=kw(J)P(X=k)=\sum_{\substack{J\subseteq\{1,\ldots,n\}\\|J|=k}}w(J)
Moments
E(X)=∑ipiV(X)=∑ipiqiE(X)=\sum_i p_i\qquad V(X)=\sum_i p_iq_i
I1I_1
I2I_2
I3I_3
I4I_4
q1p2q3p4\textcolor{#ee9daf}{q_1}\textcolor{#8bded6}{p_2}\textcolor{#ee9daf}{q_3}\textcolor{#8bded6}{p_4}
one success set Jsum every pattern with k successes
Each pattern has its own product weight

Meaning

Let independent IiI_i be 1 with probability pip_i and 0 otherwise, with 0≤pi≤10\le p_i\le1. X=∑iIiX=\sum_i I_i counts the successes.

When to use it

Use for independent trials with different success chances. Each JJ names a success pattern. Add the weights for all patterns with kk successes.

Ordinary binomial requires equal pip_i. Their average preserves the mean but generally gives wrong exact-count probabilities.
Discrete count models25

Binomial distribution

Count of successes; q=1−pq=1-p
X∼Binomial⁡(n,p)k=0,1,…,nX\sim\operatorname{Binomial}(n,p)\qquad k=0,1,\ldots,n
Probability mass function
f(k)=P(X=k)=(nk)pkqn−k\textcolor{#f4dc88}{f(k)}=P(X=k)=\binom{n}{k}\textcolor{#8bded6}{p^k}\textcolor{#ee9daf}{q^{n-k}}
Moments
μX=E(X)=npσX2=V(X)=npq\mu_X=E(X)=np\qquad\sigma_X^2=V(X)=npq
pk qn−k\textcolor{#8bded6}{p^k}\,\textcolor{#ee9daf}{q^{n-k}}
00
kk
nn
choose success positions × one-pattern probability
A fixed number of independent, common-p trials

Meaning

(nk)\binom{n}{k} chooses the success positions. The powers give the probability of one arrangement with kk successes and n−kn-k failures.

When to use it

Use when there are a fixed nn binary trials, mutual independence, and the same success probability pp at every trial.

(nk)=n!k!(n−k)!\binom{n}{k}=\dfrac{n!}{k!(n-k)!} · “Counts successes” alone does not establish the model.
Discrete count models26

Translating count questions

Exactly k
P(X=k)=f(k)P(X=k)=\textcolor{#f4dc88}{f(k)}
At most k
P(X≤k)=∑j≤kf(j)P(X\le k)=\sum_{j\le k}f(j)
At least k
P(X≥k)=1−P(X≤k−1)P(X\ge k)=1-P(X\le k-1)
xx
f(j)f(j)
kk
Exactly k → at most k → at least k

Meaning

For integer counts: “fewer than kk” means X≤k−1X\le k-1. “More than kk” means X≥k+1X\ge k+1. “At least one” is the complement of zero.

When to use it

Use the chosen model’s PMF inside the sum. Limit the summation to its valid support, then use a complement if that is simpler.

P(a≤X≤b)=F(b)−F(a−1)P(a\le X\le b)=F(b)-F(a-1) for integers a≤ba\le b.
Discrete count models27

Hypergeometric distribution

Sample without replacement
X∼Hyp.Geo.⁡(N,K,n)X\sim\operatorname{Hyp.Geo.}(N,K,n)
Probability mass function
P(X=k)=(Kk)(N−Kn−k)(Nn)P(X=k)=\frac{\textcolor{#8bded6}{\binom{K}{k}}\textcolor{#ee9daf}{\binom{N-K}{n-k}}}{\binom{N}{n}}
Moments; p=K/Np=K/N and q=1−pq=1-p
μX=npσX2=npq N−nN−1\mu_X=np\qquad\sigma_X^2=npq\,\textcolor{#f4dc88}{\frac{N-n}{N-1}}
KK
N−KN-K
kk
n−kn-k
population → sample without replacement
Choose successes and failures from their own groups

Meaning

NN objects contain KK successes. A uniformly random sample of nn distinct objects gives XX successes. The draws are generally dependent.

When to use it

Use for a fixed-size sample without replacement from a known finite population. The numerator chooses both successes and failures.

max⁡(0,n−(N−K))≤k≤min⁡(n,K)\max(0,n-(N-K))\le k\le\min(n,K) · Variance formula assumes N>1N>1.
Discrete count models28

Poisson distribution

Count in a specified exposure; λ>0\lambda>0
X∼Poisson⁡(λ)k=0,1,2,…X\sim\operatorname{Poisson}(\textcolor{#f4dc88}{\lambda})\qquad k=0,1,2,\ldots
Probability mass function
f(k)=P(X=k)=e−λλkk!f(k)=P(X=k)=e^{-\lambda}\frac{\lambda^k}{k!}
Moments and exposure
μX=E(X)=σX2=V(X)=λλ=rt\mu_X=E(X)=\sigma_X^2=V(X)=\lambda\qquad\lambda=rt
λ=rt\lambda=rt
count events in this exposure
tt
E(X)=V(X)=λE(X)=V(X)=\lambda
Rate × exposure gives the expected count

Meaning

λ\lambda is the expected count for the stated exposure. Under a homogeneous Poisson process, rr is the constant rate and tt is the exposure length.

When to use it

Use for event counts under a model with independent disjoint increments and rare single events in very small intervals.

Match rate and exposure units. An average rate by itself does not justify a Poisson model.
Discrete count models29

Choosing a discrete model

Bernoulli: one binary trial
p(q=1−p)p\quad(q=1-p)
Binomial: fixed n, independent, common p
n,pn,p
Hypergeometric: n distinct draws without replacement
N,K,nN,K,n
Poisson: count over a stated exposure
λ=expected count\lambda=\text{expected count}
WHAT IS BEING COUNTED?one trialindependentcommon-p trialsfinite sampleno replacementevents duringan exposure
Start with the mechanism, then select the distribution

Meaning

The sampling mechanism determines the distribution. The desired event then determines a PMF term, sum, or complement.

When to use it

Identify what XX counts, its support, the parameters, and which assumptions the wording actually supplies.

Approximations: hypergeometric ≈\approx binomial for a small sampling fraction; binomial ≈\approx Poisson for large nn, small pp.
Discrete count models30

A reliable setup and answer check

1

Identify the object

A statistic, an event, or a random variable

2

Translate the wording

The reference group, support, and event endpoints

3

Justify the formula

The sampling mechanism and required assumptions

4

Interpret and check

Probability bounds, units, signs, and plausibility

1object2event3assumptions4answer
A setup that explains why the formula applies
Chapters 1–531