Back
Keentune
Statistics curriculum 9 chapters
·
59 concepts
·
free
Everything the adaptive question bank can teach and test in Statistics, from foundations through advanced practice. Work through it in order, or start practicing and let the questions find your level.
Start practicing Statistics
New here? Read the Statistics guide A free 8-minute primer — the mental model, the mistakes beginners make, and what to practice first.
A. Measures of center •
Computing the mean as the sum divided by the count.
The mean adds the values and divides by how many there are: 37, 52, 46, 61, and 44 sum to 240, so the mean is 240 ÷ 5 = 48. The mean must land between the smallest and largest values, a quick check on the arithmetic. Dividing by the wrong count is the most common slip.
•
Sorting first, then taking the middle value.
The median is the middle value once the data are sorted. For 73, 58, 91, 64, and 80, sort to 58, 64, 73, 80, 91; the middle of five values is the third, 73. Taking the middle of the unsorted list is the classic error.
•
Averaging the two middle values when the count is even.
With an even number of values there are two middle values, and the median is their average. Sorting 105, 87, 96, and 112 gives 87, 96, 105, 112, so the median is (96 + 105) ÷ 2 = 100.5. The median of an even-sized set need not be one of the data values.
•
Finding the most frequent value.
The mode is the value that occurs most often: in 7, 11, 7, 15, 11, 7, the value 7 appears three times, so it is the mode. A set can have no mode, when every value appears once, or several, when two or more values tie for the most appearances.
•
Averaging the largest and smallest values.
The midrange is the average of the largest and smallest values: for 83, 47, 65, and 71, it is (83 + 47) ÷ 2 = 65. It uses only the two extremes, so a single unusual value moves it as much as that value moves the extreme.
•
Using the mode as the only center for categorical data.
For categories such as blood type or car color, the values cannot be added or ordered, so the mean and median make no sense; the mode is the only measure of center. Among blood types A, O, B, O, AB, and O, the typical type is O.
•
Finding a missing value through the total a mean implies.
A mean times its count gives the total. If 6 numbers have a mean of 14, they total 84; if five of them add to 71, the sixth is 84 − 71 = 13. Working through the total is faster and safer than guessing values that average to 14.
•
Updating a mean when a value is added.
When a value joins a set, add it to the old total and divide by the new count. Seven numbers averaging 32 total 224; adding 56 gives 280 over 8 values, a mean of 35. Averaging the old mean with the new value, which gives 44, weights the newcomer as heavily as all seven others.
•
Weighting each value by its share of the total.
A weighted mean multiplies each value by its weight and adds the results. A grade that is 55% exam and 45% homework, with scores of 80 and 60, is 0.55 × 80 + 0.45 × 60 = 44 + 27 = 71. The plain average, 70, ignores that the exam counts for more.
•
Combining group means through their totals.
Means of groups with different sizes combine through their totals. A class of 12 averaging 74 and one of 36 averaging 82 total 888 and 2,952, so all 48 students average 3,840 ÷ 48 = 80, not the midpoint 78. The larger group pulls the combined mean toward its own.
Practice this section →
B. Measures of spread •
Computing the range as the largest minus the smallest value.
The range is the largest value minus the smallest: for 14, 39, 22, 51, and 30, it is 51 − 14 = 37. It is quick but depends on only two values, so a single extreme value can make a tight set look widely spread.
•
Comparing how spread out two sets are, not where they are centered.
Two sets can share a center and differ in spread. The sets 69, 70, 71 and 40, 70, 100 both have mean 70, but the first ranges over 2 and the second over 60, so the second is far more spread out. A question about spread is not answered by comparing the means.
•
Computing the population variance as the mean squared deviation.
The population variance is the average squared distance from the mean. For 3, 7, 7, and 11 the mean is 7, the deviations are −4, 0, 0, and 4, their squares are 16, 0, 0, and 16, and the variance is 32 ÷ 4 = 8. Squaring keeps negative and positive deviations from canceling.
•
Taking the square root of the variance.
The standard deviation is the square root of the variance, which puts spread back in the data's own units. For 5, 11, 11, and 13 the mean is 10, the squared deviations are 25, 1, 1, and 9, the variance is 36 ÷ 4 = 9, and the standard deviation is 3. Forgetting the square root reports the variance instead.
•
Recognizing that identical values have no spread.
If every value in a set is the same, every deviation from the mean is zero, so the range, the variance, and the standard deviation are all 0. Spread measures can never be negative, which is a quick check on a computed answer.
•
Keeping the units of each spread measure straight.
The range and the standard deviation are in the data's units, while the variance is in squared units: heights in centimeters have a variance in square centimeters. Comparing a standard deviation directly with a variance mixes units and leads to wrong conclusions.
•
Dividing by n for a population and by n − 1 for a sample.
A population variance divides the sum of squared deviations by n, the number of values. A sample variance, used to estimate a larger population, divides by n − 1, which corrects for the sample's tendency to understate spread. Read whether a question asks for the population version.
Practice this section →
C. Outliers and shape •
Seeing how one extreme value moves the mean but not the median.
One extreme value can drag the mean far from the typical values while the median barely moves. For 31, 28, 35, 30, and 410, the mean is 534 ÷ 5 = 106.8, while the median is 31. The mean is pulled toward the outlier, and in that direction.
•
Seeing a low outlier pull the mean down.
An unusually small value pulls the mean down just as a large one pulls it up. For 62, 58, 3, 61, and 66, the mean is 250 ÷ 5 = 50, while the median is 61. The direction of the pull matches the side the outlier is on.
•
Using the median as the typical value when data are skewed.
When a few values are far larger than the rest, as with commute times when a few trips take hours, the mean overstates the typical value and the median describes it better. The median depends only on the middle of the sorted data, so the extremes do not move it.
•
Knowing which measures resist extreme values.
The median and the interquartile range, the spread of the middle half of the data, resist extreme values; the mean, the range, and the standard deviation do not. When a data set contains an outlier, a resistant measure describes the bulk of the data, while a sensitive one mostly reports the outlier.
Practice this section →
D. Populations and samples •
Telling the population from the sample.
The population is every individual you want to draw conclusions about; a sample is the part you actually measure. A study of 400 students drawn from a district of 30,000 measures a sample in order to learn about the district's population.
•
Telling a parameter from a statistic.
A parameter is a number that describes a population, such as the true mean height of all adults in a country, and is usually unknown. A statistic is a number computed from a sample, such as the mean height of 500 measured adults, and is used to estimate the parameter.
•
Telling a census from a sample survey.
A census measures every member of the population, while a sample survey measures only some and generalizes. A census removes sampling error but is costly and slow, which is why most studies use carefully chosen samples instead.
•
Choosing a sample by chance so that it represents the population.
In a random sample every member of the population has a known chance of selection, so no group is favored by how the sample was chosen. Randomness protects against bias in selection; it does not guarantee that any single sample is perfectly representative.
•
Sampling randomly within each subgroup.
Stratified sampling divides the population into groups, such as age bands or regions, and draws a random sample from each. It guarantees that every group is represented, and it often gives more precise estimates than one simple random sample of the same size.
•
Seeing that larger samples give less variable estimates.
Estimates from larger random samples vary less from sample to sample, so they tend to land closer to the true value. Size cannot fix a biased method, though: a huge sample drawn the wrong way can still be far off.
•
Understanding how a statistic varies across samples.
If you drew many random samples and computed the same statistic on each, its values would form a sampling distribution. The standard error is that distribution's standard deviation: the typical distance between a sample estimate and the true value. Larger samples give a smaller standard error.
Practice this section →
E. Bias in data collection •
Recognizing a sample chosen in a way that misrepresents the population.
Selection bias occurs when the way a sample is chosen makes it systematically different from the population. Asking only people at a gym how often adults exercise overstates exercise, because the sample was drawn from where exercisers are. A convenience sample, of whoever is easiest to reach, is the common cause.
•
Recognizing distortion when nonresponders differ from responders.
Nonresponse bias arises when the people who do not answer differ from those who do. A phone survey on working hours that busy workers rarely answer will understate working hours, however carefully the numbers were dialed.
•
Recognizing bias when people choose to respond.
When people choose whether to respond, as in a call-in radio vote, those with strong opinions are overrepresented, and the result says little about the population. Voluntary response samples are biased even when they are very large.
•
Recognizing conclusions drawn only from survivors.
Survivorship bias comes from studying only the cases that made it through some filter. Studying only companies still in business to learn what makes firms last ignores the firms that did the same things and failed, so the lessons can be illusions.
•
Recognizing questions that push respondents toward an answer.
Loaded or leading wording biases a survey: "Do you support the reckless plan to close the library?" invites a no before the question is weighed. Neutral wording, such as "Do you support or oppose closing the library?", lets the answers reflect opinion rather than phrasing.
Practice this section →
F. Study design •
Telling an observational study from an experiment.
An observational study measures people as they are, without assigning anything; an experiment assigns a treatment and compares outcomes. Only a well-designed experiment can support a claim that the treatment caused the difference.
•
Using chance to decide who gets each treatment.
Random assignment uses chance to decide who receives each treatment, so the groups tend to be alike in everything except the treatment, including traits nobody thought to measure. That is what lets an experiment attribute a difference in outcomes to the treatment.
•
Using a placebo and blinding to remove expectation effects.
A placebo is an inactive treatment that looks like the real one, so the comparison group has the same experience apart from the active ingredient. In a double-blind study neither the participants nor the people measuring outcomes know who received which, so expectations cannot sway the results.
•
Recognizing a third variable that drives both sides of an association.
A confounding variable is linked to both the supposed cause and the effect and can create an association between them. Towns with more fire stations report more fires, but town size drives both; adding fire stations does not cause fires.
•
Considering that the effect may run the other way.
An association can run in the opposite direction from the one claimed. People who take a pain reliever report more pain than people who do not, but the pain leads to the medicine, not the reverse. Ask which variable could plausibly come first.
Practice this section →
G. Correlation •
Reading whether two variables move together or in opposite directions.
Two variables are positively associated when high values of one go with high values of the other, and negatively associated when high values of one go with low values of the other. Air temperature falls as altitude rises, a negative association.
•
Reading the strength and sign of r.
The correlation coefficient r runs from −1 to 1. Its sign gives the direction, and its distance from 0 gives the strength of a straight-line relationship: r = 0.92 is a strong positive association, while r = −0.3 is a weak negative one.
•
Interpreting a correlation near zero.
A correlation near 0 means there is no straight-line relationship, not necessarily no relationship at all: points along a U-shaped curve can have r close to 0. It also says nothing about cause.
•
Refusing to read cause from correlation alone.
A strong correlation does not show that one variable causes the other. A confounder, reverse causation, or chance can produce the same pattern, and only a well-designed experiment, or a careful account of those alternatives, supports a causal claim.
•
Avoiding predictions far outside the observed data.
A trend line describes the data it was fitted to and may not hold beyond them. A line fitted to children's heights from ages 2 to 10 would predict absurd heights at 40, because growth stops. Predict only within, or close to, the range of the observed data.
Practice this section →
H. Hypothesis tests and intervals •
Setting up the null and alternative hypotheses.
A test begins with two competing claims. The null hypothesis says nothing is going on: no effect, and no difference between groups. The alternative is what the study hopes to show. The test asks whether the data would be surprising enough, if the null held, to reject it in favor of the alternative.
•
Interpreting a p-value.
A p-value answers one question: if the null hypothesis held, how often would a result at least this far from it turn up? A tiny p-value says such data would rarely occur under the null. It does not give the probability that the null hypothesis is true, which is the most common misreading.
•
Comparing a p-value with the significance level set in advance.
The significance level is a threshold chosen before the test, often 0.05. If the p-value falls below it, the result is called statistically significant and the null is rejected. Setting the level at 0.05 means accepting a 5% chance of rejecting a true null hypothesis.
•
Telling a false positive from a false negative.
A Type I error is a false alarm: the test rejects the null even though the null is correct. A Type II error is a miss: the null is wrong, but the test fails to reject it. Lowering the significance level makes false alarms rarer and, other things equal, misses more common.
•
Understanding a test's power to detect a real effect.
The power of a test is its probability of rejecting the null hypothesis when a real effect exists. Larger samples, larger true effects, and less noisy measurements all raise power. A test with low power often misses effects that are really there.
•
Reading a non-significant result correctly.
Failing to reject the null hypothesis means the data were not strong enough evidence against it, not that the null has been proved. A small or noisy study can easily miss a real effect, so absence of evidence is not evidence of absence.
•
Separating statistical significance from practical importance.
A statistically significant result can still be too small to matter. With a very large sample, a difference of half a point on a 100-point test can reach significance while being useless in practice. Always ask how big the effect is, not only whether it is significant.
•
Recognizing false positives from running many tests.
Each test of a true null hypothesis at the 0.05 level carries a 5% chance of a false alarm, so a study that runs dozens of tests should expect a few spurious hits. Reporting only the comparisons that came out significant, sometimes called p-hacking, presents those chance results as discoveries.
•
Interpreting a confidence interval.
A 95% confidence interval comes from a method that captures the true parameter in about 95% of repeated samples. A single interval either contains the true value or does not; the 95% describes how reliable the method is, not the chance for that particular interval.
•
Seeing how sample size sets the margin of error.
The margin of error is the half-width of a confidence interval: the plus-or-minus around an estimate. Larger samples make it smaller, because estimates from them vary less. Quadrupling the sample size roughly halves the margin of error.
Practice this section →
I. Common pitfalls •
Weighing a test result against how common the condition is.
A positive result from an accurate test can still be more likely false than true when the condition is rare. Take a condition affecting 1 in 500 people and a test that catches 98% of cases but falsely flags 3% of everyone else. Among 50,000 people, 98 true positives sit beside 1,497 false ones, so only about 6% of positives are real.
•
Expecting extreme results to be followed by more typical ones.
An extreme result is partly luck, so the next measurement tends to be closer to average. Patients often join a trial when their symptoms are at their worst, so many improve afterward even with no treatment at all. Mistaking that drift back toward the average for the effect of an intervention is a common error.
•
Expecting small groups to produce the most extreme rates.
Small groups produce more extreme averages and rates simply because chance has more effect on them. The counties with the highest and the lowest rates of a rare disease are often the smallest counties, not counties that are especially healthy or unhealthy.
•
Recognizing a trend that reverses when groups are combined.
Simpson's paradox occurs when a pattern that holds within every subgroup reverses once the subgroups are combined. It happens when the groups differ in size and in a lurking variable, so comparisons should be made within comparable groups before pooling.
•
Reading a chart's axis before judging the size of a difference.
A chart whose vertical axis starts well above zero exaggerates differences: with an axis starting at 40, bars of 42 and 44 look like one is double the other. Read the axis labels before judging how large a difference really is.
•
Recognizing selective use of data.
Cherry-picking reports only the data that support a conclusion, such as the one month when sales rose, and leaves out the rest. Ask what the full data set shows and how the reported part was chosen.
Practice this section →
Keentune is not affiliated with or endorsed by the organizations whose documentation informs these maps.
Start practicing Statistics
All about Statistics practice
Browse all free study guides
All exam, test, and product names and trademarks are the property of their respective owners and are used here for identification and reference only. Keentune is independent study practice — not affiliated with, authorized, or endorsed by any of these organizations.
© 2026 SportaApp LLC