📖 Crammy · All study guides
AP Statistics · Unit 2

Exploring Bivariate Data: every key term you need

22 flashcard terms for AP Statistics Unit 2, written to match the course framework. Study them here, then drill them as interactive flashcards — free, no account needed.

Study this unit free →
Exploring Bivariate Data
Examining relationship between two variables; key questions: Is there association? How strong? What's the pattern?
Scatterplot
Plot with one variable on x-axis, other on y-axis; shows association pattern (linear, curved, clustered, none).
Correlation Coefficient (r)
Measures strength/direction of LINEAR relationship: r=-1 (perfect negative), r=0 (no linear), r=+1 (perfect positive).
Correlation Properties
Unitless, ranges -1 to +1. Same r if x/y swapped. r ≠ causation. Sensitive to outliers. Only measures linear relationships.
Causation vs Correlation
Strong correlation ≠ causation. May have: reverse causation, confounding variables, lurking variables, coincidence.
Least Squares Regression Line
Line minimizing squared vertical distances (residuals) from points. Equation: ŷ = a + bx where b=r(sy/sx).
Regression Line Interpretation
Slope b: for each 1-unit increase in x, y increases by b units (on average). Intercept a: predicted y when x=0.
Predictions & Extrapolation
Use regression to predict y for given x (interpolation is OK). Extrapolation (beyond data range) unreliable.
Residuals
Actual y - Predicted ŷ. Residual plot (residuals vs x) should show random scatter (no pattern) if linear model appropriate.
Coefficient of Determination (R²)
R² = r²; proportion of y's variation explained by x. R²=0.85 means 85% of variation in y explained by x.
Outliers & Influential Points
Outlier: unusual y value. Influential: point far from others on x-axis, dramatically changes regression line.
Transformations
When linear model fails, transform data (log, square root) to achieve linearity; analyze transformed data.
Categorical Variables
Use indicator (dummy) variable (0/1) as predictor. Regression slope compares mean outcome between categories.
Simpson's Paradox
Trend reverses when data subdivided by group. Shows importance of examining relationships within subgroups.
Ecological Fallacy
Drawing individual-level conclusions from group-level data (e.g., states' data doesn't apply to individuals in states).
Spurious Correlation
Strong correlation exists due to confounding variable, not direct causation. Classic example: ice cream sales & drownings (temperature).
Two-Way Tables
Frequency table for two categorical variables. Examine marginal (row/column) distributions and conditional distributions.
Conditional Distribution
Distribution of one variable given specific value of another. Shows how relationship differs within subgroups.
Independence
Two categorical variables independent if conditional distributions same as marginal distribution (no association).
Chi-Square Association
Measures strength of association between categorical variables. Higher value = stronger association.
Mosaic Plots
Visual display of two-way table; tile size represents cell frequency; useful for categorical relationships.
Unit 2 Key Ideas
Correlation measures linear association; r close to ±1 = strong, r≈0 = weak. Regression predicts y from x but doesn't imply causation. Residual plots reveal if linear model appropriate.
Turn these into flashcards & quizzes →

More AP Statistics guides