Statistics questions mix computation with judgment. Some ask for a mean, a median, a mode, a range, or a standard deviation; others ask what a study can conclude, whether a sample is biased, or what a p-value means. Most wrong answers come from a skipped step, such as taking the median of unsorted data, or from reading more into the data than the design allows. The chapters cover the computations, the ideas behind sampling and study design, and the reasoning behind tests and intervals.
Each chapter opens with the short version. Tap one to read the detail.
Center and outliers
~2 min
The mean is the sum over the count, the median is the middle of the sorted data, and the mode is the most frequent value. An outlier drags the mean toward it while the median holds still, so skewed data are better summarized by the median.
The mean adds the values and divides by the count: 18, 27, 33, and 42 sum to 120, so the mean is 30. The median is the middle value once the data are sorted, and with an even count it is the average of the two middle values, so the median of 9, 4, 15, and 12 is (9 + 12) ÷ 2 = 10.5. The mode is the most frequent value, and it is the only center that works for categories such as favorite sport. The midrange averages the largest and smallest values.
Many questions are really about totals. If 5 test scores average 72 and four of them add up to 290, the fifth is 360 − 290 = 70. A weighted mean multiplies each value by its weight: a portfolio with 30% in a fund that returned 10% and 70% in one that returned 4% earned 0.3 × 10 + 0.7 × 4 = 5.8%, not the plain average of 7%.
An outlier pulls the mean toward itself while barely moving the median. For 44, 47, 41, 45, and 213, the mean is 78 but the median is 45, so the median describes the typical value far better. When a few very large or very small values skew the data, report the median.
Rule: sort before taking a median, work averages through their totals, and prefer the median when a few extreme values skew the data.
Measuring spread
~2 min
The range is the largest value minus the smallest. The variance averages the squared distances from the mean, and the standard deviation is its square root, in the data's own units.
The range is the simplest measure of spread: for 23, 8, 41, 17, and 30, it is 41 − 8 = 33. It depends on only two values, so one extreme value can inflate it. Two sets can share a center and still differ in spread: 49, 50, 51 and 20, 50, 80 both average 50, but the second is far more spread out.
The population variance averages the squared distances from the mean. For 6, 10, 14, and 18, the mean is 12, the deviations are −6, −2, 2, and 6, their squares add up to 80, and the variance is 80 ÷ 4 = 20. The standard deviation is the square root of the variance, about 4.47 here. It is in the data's own units, while the variance is in squared units. A sample variance, used to estimate a larger population, divides by n − 1 instead of n.
Rule: square the deviations before averaging them, take the square root for the standard deviation, and remember that one extreme value can inflate the range.
Samples and bias
~2 min
A sample stands in for a population, a statistic estimates a parameter, and random selection keeps the sample representative. Bias enters through how a sample is chosen, who responds, and how questions are worded, and a bigger sample does not cure it.
A population is everyone a study wants to describe, and a sample is the part it actually measures. A number computed from a sample, called a statistic, estimates the matching number for the population, called a parameter. Random sampling gives every member a known chance of selection, and stratified sampling draws randomly within groups, such as regions, so that each group is represented.
Larger random samples give estimates that vary less from sample to sample, but size cannot fix a biased method. Selection bias comes from where the sample is drawn: surveying travelers at an airport about how often people fly overstates flying. Nonresponse bias appears when the people who answer differ from those who do not, and voluntary response, where people choose to take part, overrepresents strong opinions. Survivorship bias studies only the cases that passed some filter, and leading wording pushes respondents toward an answer.
Rule: judge a sample by how it was chosen and who answered, not by its size alone.
Study design and correlation
~2 min
Only a randomized experiment can support a cause-and-effect claim; an observational study can show an association. Correlation measures the strength and direction of a straight-line association, and confounders, reverse causation, and extrapolation are the usual traps.
An observational study measures people as they are, while an experiment assigns treatments. Random assignment makes the groups alike in everything but the treatment, so a difference in outcomes can be traced to the treatment. A placebo gives the comparison group the same experience without the active ingredient, and in a double-blind study neither the participants nor the evaluators know who got what.
The correlation coefficient r runs from −1 to 1. Its sign gives the direction of a straight-line association and its size the strength, so r = 0.35 is a fairly weak positive association. A value near 0 rules out only a straight-line pattern. Correlation alone does not show cause. A confounder can drive both variables: cities with more bookstores report more burglaries, but population drives both. The effect can also run backward, from the supposed result to the supposed cause. And a trend line fitted to one range of data can fail badly when it is extended far beyond that range.
Rule: credit a cause only to a randomized experiment, and before trusting a correlation, look for a confounder, reverse causation, or an extrapolation.
Tests and intervals
~2 min
A test asks whether the data are surprising under a null hypothesis; the p-value measures that surprise and is compared with a significance level chosen in advance. A confidence interval comes from a method that captures the true value at a stated rate, and larger samples make it narrower.
A hypothesis test starts from a null hypothesis of no effect and asks how surprising the data would be if it were true. The p-value is the chance, if the null were true, of a result at least as extreme as the one observed; it is not the chance that the null is true. If the p-value falls below a significance level fixed in advance, commonly 0.05, the null is rejected, and that level is also the rate of false positives the test accepts when the null is true.
A Type I error rejects a true null, and a Type II error fails to reject a false one. Power, the chance of detecting a real effect, rises with the sample size and the size of the effect. A non-significant result means the evidence was too weak, not that the null was proved. A significant result may still be too small to matter, and running many tests produces some false positives by chance, so a lone significant result among many deserves suspicion.
A 95% confidence interval comes from a method that captures the true value in about 95% of repeated samples. Its half-width is the margin of error, which shrinks as the sample grows: quadrupling the sample roughly halves it.
Rule: read a p-value as the surprise of the data under the null, never as the chance that the null is true, and read an interval as the output of a method with a stated success rate.
Common pitfalls
~2 min
Weigh a test result against how rare the condition is, expect extreme results to drift back toward average, expect small groups to produce extreme rates, and read a chart's axis before judging a difference.
A positive result from an accurate test can still be more likely false than true when the condition is rare. Suppose 1 in 250 people have a condition, the test catches 95% of them, and it wrongly flags 4% of everyone else. Out of 25,000 people, 100 have the condition and 95 of them test positive, while 996 of the 24,900 without it also test positive, so only about 9% of the positives are real.
Extreme results are partly luck, so a second measurement tends to sit closer to average; a drop after a record month may be regression to the mean rather than a real decline. Small groups produce more extreme rates simply by chance, so the top and bottom of a ranking are often filled with small groups. Simpson's paradox can reverse a trend when groups of different sizes are combined, so compare within comparable groups first.
Charts mislead too: a vertical axis that starts well above zero makes small differences look large, and a report that shows only the favorable months hides the rest of the data.
Rule: before trusting a striking result, check the base rate, the sample size, the comparison groups, and the axis.
Written by Keentune. We are not affiliated with or endorsed by the organizations whose documentation informs this guide, and any linked sources belong to their respective owners.
All exam, test, and product names and trademarks are the property of their respective owners and are used here for identification and reference only. Keentune is independent study practice — not affiliated with, authorized, or endorsed by any of these organizations.