๐Ÿ“– Crammy ยท All study guides
Probability & Statistical Inference ยท Topic 8

Hypothesis Testing, Errors and Power: every key term you need (+ practice quiz)

25 flashcard terms for Probability & Statistical Inference Topic 8, written to match the course framework. Study them here, then drill them as interactive flashcards, or test yourself with the 15-question quiz โ€” free, no account needed.

Study this unit free โ†’
Null hypothesis
The default claim about a parameter that the data are given a chance to contradict, usually a statement of no effect or no difference. It is never proved, only retained for lack of evidence against it.
Alternative hypothesis
The claim the analyst entertains if the null is rejected, stated before the data are seen. Its direction determines whether the test is one sided or two sided and therefore how the tail area is computed.
Test statistic
A number computed from data that measures how far the sample sits from what the null predicts, calibrated in standard error units. Its null distribution is what turns the number into a probability statement.
Rejection region
The set of test statistic values leading to rejection of the null, chosen so that the probability of landing there under the null equals the significance level. It is fixed before the data arrive.
Significance level
The pre-chosen probability of rejecting a true null hypothesis, commonly set at five percent. It is a tolerance for false alarms selected by the analyst, not a fact about the data or the science.
P-value
The probability, computed assuming the null is true, of a test statistic at least as extreme as the one observed. It measures compatibility of data with the null, not the probability that the null is true.
P-value misinterpretation
A p-value is not the probability of the null hypothesis, not the probability the result is a fluke, and not a measure of effect size. Small values indicate surprise under the null, nothing more.
Type I error
Rejecting a null hypothesis that is actually true, a false positive. Its probability is controlled directly by the significance level whenever the test's assumptions hold.
Type II error
Failing to reject a null hypothesis that is actually false, a false negative. Its probability depends on the true effect size, the spread of the data and the sample size, so it is not chosen directly.
Power of a test
The probability of rejecting the null when a specified alternative is true, equal to one minus the type II error rate. It rises with sample size, effect size and lower measurement noise.
Power analysis
Calculating the sample size needed to detect an effect of a specified size with a target power. Done before data collection it prevents wasted studies; done afterwards from the observed effect it is uninformative.
Effect size measure
A standardised description of how large a difference is, independent of sample size. It answers the question of practical importance that a p-value cannot address on its own.
Statistical versus practical significance
A large sample can make a trivial difference statistically significant, while a small sample can leave an important difference undetected. Significance and importance are separate judgements and must be reported separately.
One-sided versus two-sided tests
A one-sided test concentrates the rejection region in a single tail and gains power there while abandoning the other direction entirely. Choosing the side after seeing the data invalidates the stated error rate.
Failure to reject
The correct description of a nonsignificant result, meaning the evidence was insufficient rather than the null being demonstrated. Absence of evidence with low power is very weak evidence of absence.
Neyman-Pearson lemma
For a simple null against a simple alternative, the likelihood ratio test has the greatest power among all tests with the same size. It is the theoretical justification for likelihood-based test statistics.
Likelihood ratio test
Comparing the maximised likelihood under the restricted null with that under the unrestricted model. Twice the log of the ratio is approximately chi-square in large samples with degrees of freedom equal to the number of constraints.
Wald and score tests
Two large-sample alternatives to the likelihood ratio, one based on the distance of the estimate from the null scaled by information, the other on the score evaluated at the null. All three agree asymptotically but can differ in finite samples.
Uniformly most powerful test
A test that maximises power against every alternative in a family at a fixed size. Such tests exist for one-sided hypotheses in many standard families but generally not for two-sided ones.
Multiple comparisons problem
Running many tests at the same level makes at least one false rejection nearly certain even when every null is true. Reported findings must be adjusted or pre-registered to keep the overall error rate meaningful.
Bonferroni correction
Dividing the significance level by the number of tests to control the chance of any false rejection. It is simple and always valid but conservative, sacrificing power when tests are numerous or correlated.
False discovery rate control
Limiting the expected share of rejected hypotheses that are false positives instead of the chance of any error at all. It is far more powerful than family-wise control in large screening studies.
Data dredging
Searching many models, subgroups or outcomes and reporting only what reached significance. The stated error rate then describes a procedure that was never followed, and published effects are systematically overstated.
Assumption violations in testing
Dependence between observations, unequal variances or heavy tails distort the null distribution, so the nominal error rate is wrong. Dependence is the most damaging because it silently shrinks the effective sample size.
Permutation test
Building a null distribution by repeatedly reshuffling group labels, valid under an exchangeability assumption rather than a parametric shape. It is a natural fallback when standard distributional assumptions look untenable.
Turn these into flashcards & quizzes โ†’

More Probability & Statistical Inference guides