Expectation, Variance and Moment Generating Functions: every key term you need (+ practice quiz)
25 flashcard terms for Probability & Statistical Inference Topic 3, written to match the course framework. Study them here, then drill them as interactive flashcards, or test yourself with the 15-question quiz โ free, no account needed.
The probability-weighted average of a random variable, a sum for discrete cases and an integral for continuous ones. It is a long-run centre of mass and need not be a value the variable can actually take.
Law of the unconscious statistician
The expectation of a function of a random variable can be computed by averaging that function against the original mass or density, without first deriving the distribution of the transformed variable.
Linearity of expectation
The expectation of a sum equals the sum of expectations, with constants factoring out. This holds whether or not the variables are independent, which makes it the most robust tool in the whole subject.
Nonexistent expectation
Some distributions, such as the Cauchy, have integrals that fail to converge absolutely, so no mean exists. Sample averages from such models wander instead of settling, and standard limit theorems do not apply.
Variance
The expected squared deviation from the mean, measuring spread in squared units. It equals the mean of the square minus the square of the mean, the identity used in almost every hand computation.
Standard deviation
The square root of the variance, restoring the original units of measurement. It is the natural scale for describing typical deviation and for standardising a variable before comparison.
Variance of a linear transformation
Adding a constant leaves variance unchanged while multiplying by a constant multiplies variance by that constant squared. Shifts move location only; scaling stretches spread quadratically.
Standardisation
Subtracting the mean and dividing by the standard deviation to produce a variable with mean zero and unit variance. It makes distributions comparable and is the step behind every table of critical values.
Moment about the origin
The expectation of a whole-number power of a random variable. The first such moment is the mean, and higher ones feed into descriptions of spread, asymmetry and tail weight.
Central moment
The expectation of a power of the deviation from the mean. The second is the variance, and the third and fourth are the raw material for skewness and kurtosis measures.
Skewness
A standardised third central moment measuring asymmetry. Positive values indicate a long right tail, which warns that the mean sits above the median and that symmetric approximations may be poor in small samples.
Kurtosis
A standardised fourth central moment describing tail weight relative to the normal case. Heavy tails inflate the variability of sample variances and slow the convergence promised by limit theorems.
Moment generating function
The expectation of the exponential of a parameter times the variable, defined where that expectation is finite. Differentiating it at zero produces the moments in order, turning integration into differentiation.
Uniqueness property of generating functions
If two distributions have moment generating functions that agree on an open interval containing zero, the distributions are identical. This is the standard route for identifying the law of a sum.
Generating function of a sum
For independent variables the moment generating function of the sum is the product of the individual functions. This single fact establishes most of the closure results for normal, Poisson and gamma families.
Characteristic function
The expectation of a complex exponential of the variable, which always exists even when no moment generating function does. It is the rigorous tool behind proofs of the central limit theorem.
Probability generating function
For nonnegative integer variables, the expectation of a dummy variable raised to the random power. Its derivatives at one recover factorial moments and its coefficients recover the mass function.
Markov's inequality
For a nonnegative variable, the probability of exceeding a threshold is at most the mean divided by that threshold. It needs only the mean and gives crude but assumption-light tail control.
Chebyshev's inequality
The probability of lying more than k standard deviations from the mean is at most one over k squared, for any distribution with finite variance. It is loose but distribution-free, unlike normal-based rules.
Jensen's inequality
For a convex function, the expectation of the function is at least the function of the expectation. It explains why plugging a mean into a nonlinear formula generally gives a biased answer.
Conditional expectation
The mean of one variable computed within the subpopulation defined by a value of another, itself a random variable when the conditioning value varies. It is the best mean-squared predictor of the first variable.
Tower property
The expectation of a conditional expectation equals the unconditional expectation. Averaging over the conditioning variable recovers the overall mean, which lets hard expectations be computed in stages.
Law of total variance
Total variance splits into the mean of the conditional variances plus the variance of the conditional means. The decomposition separates within-group noise from between-group differences.
Expectation of a product
For independent variables the expectation of a product factors into the product of expectations. Without independence this fails, and the gap is exactly the covariance term.
Mean squared error decomposition
The average squared distance between an estimator and the target splits into its variance plus the square of its bias. This identity underlies the whole trade-off between precision and accuracy.