Story 9 — Normal distribution, Use number 1 To describe, to organize data

 

 

 

Michelangelo, Sistine Chapel

 

Point of contact.

God’s hand makes contact with the hand of

Man. Genesis. Magic moment. A whole

world begins here. The divine, the

immaterial, the perfect makes contact with

the earthly, imperfect, and imparts to it

some of the harmony of the spiritual, perfect

world.

 

 

Drama

Where does Basita fall?

 

Susan, our psychology professor,

decided to take a personal interest in the

learning of her students, and called

those scoring very low to her office.

Among those she called was Basita.

 

The bottom line of this is that you should

quit college immediately. You are the

bottom of the bottommost. You will

never be able to compete with other

college students. Find a job in a diner, in

a farm, anywhere, but do not waste your

time at college, she said to Basita.

 

The next day, Basita and her mother,

Mrs. Thinlips, an accountant by

profession, marched into Susan’s office.

 

I have already talked to your chairperson

about this. I demand that you explain to

me the basis of your criticism and absurd

advice to my daughter. You traumatized

her, in effect telling her that she is an

idiot. You will hear from my lawyer. For

now I want an explanation.

 

My daughter scored 45. The mean was

60. Forty-five is close to the mean, only

15 points below. Forty-five means that

Basita knows almost half of what you

expect her to know. Your telling my

daughter to quit college is most

unwarranted. I demand an explanation!

Mrs. Thinlips said, banging her fist on the

Susan’s desk.

 

Help, Susan said to herself, Goddess

Normal Curve, help. She brings out a

sizable cardboard model of the goddess

and bows.

 

Mrs., Thinlips, she said. The mean of the

scores in Basita’s class was indeed 60,

and the standard deviation was 5. Here

is the computer analysis.

 

Now we place 60 on the mean (0

standard deviation), that is in the middle

of the curve.

 

Flash, thunder, tempest winds,

Michelangelo hovers over the cardboard

model! Angels and ministers of heaven

and hell! Point of contact of the spiritual

with the material! A new science is born.

Statistics. All else is humble things after

the cosmogony of this moment of

Genesis. All subsequent statistical tests

bow to this archetypal creation.

We place 60 on the mean, that is in the

middle of the curve, Susan continues. Now

we move down to standard deviation -1,

to the first vertical line on the left of the

midline. This means that at this point we

have score 55. Now we move down one

more standard deviation, standard

deviation -2. Here we have score 50.

Finally, we move down one more standard

deviation, standard deviation -3. Here we

have score 45. This is Basita’s score. The

percentage of scores above this point is

99.865%. The lower-tail probability is about

0.135%, or roughly 1.35 out of 1000

students in expectation. This is not a fixed

maximum; observed counts vary from sample to

sample. Basita is still near the extreme lower tail. Imagine a line of 1000 students, a

small town, and your daughter standing at

the very end! Susan said, with a malicious

smile on her face.

Mrs. Thinlips or Basita have not been seen

on the campus ever since.

 

Back to our task to understand the

normal distribution, to understand

it our way, a gut-level

understanding.

 

In doing science we have two

domains, two worlds. The

empirical domain, the mud and

flesh domain, and the formal

domain, the domain of

abstractions, ideas, logic and

mathematics. The empirical

domain is our sense world, and

the data we get by running

experiments in it.

 

The formal domain is the world of

thought and mathematics.

Sciences progress by

superimposing perfect models of

mathematics on the imperfect,

variable, messy world of matter.

When we do that, we immediately

see things that we could not see

by looking only at the data we

have collected from observations

in the material world. 

 

Newton succeeded in creating a

revolution in Physics by first

creating a calculus, which he

superimposed on nature. Galileo

Galilei, the man who started

science as we know it today, said

that the language of nature is

mathematics.

 

A most important note in Basita’s

story:

 

What if Basita’s score was not 45

but it was 43? How would we find

where it falls on the normal curve?

There is a formula called the z

formula. Here it is:

 

 

 

 

 

Let’s try it.

 

Score 43 minus the mean, which

is 60, equals -17. Now if we divide

-17 by the standard deviation

which is 5 here, we get a z of -3.4.

That makes sense. Basita’s score 

of 45 fell exactly on standard

deviation -3, as we saw. A score of

43 will be even more to the left of

the curve.

 

I do not want to close this talk. I

want to play some triumphant

march, Beethoven’s Eroica

perhaps. Look at this formula. Play

with it, do things with it. Digest

what we do with it. Let’s dramatize

this.
 

 

 

Drama

An archetypal ceremony

 

I pick a score, and wave it in the

air. Then I wear my glasses and

stick my nose on the normal

curve, running up and down the

line with standard deviations on

it, I mumble:

 

Where does this score fall? Where

does this score fall?

 

I then use the z formula and find

where exactly my score falls.

 

This is an archetypal ceremony.

Remember it. We will act it out

again in the future.

 

Story 8 — The uses of the normal distribution

 

You ask:

 

Why learn all these things about

the normal curve? What is the use

of all of this?

 

The use of all of this is necessary,

I say.

 

The normal curve (the standard normal curve) is a mathematical, perfect curve, 

With magic qualities and powers.

Understanding the normal curve is

necessary, if we wish to

understand statistics from simple

t-tests to complex Analysis of

Variance (ANOVA). 

 

The normal curve is used in four

instances:

 

1. To describe, to organize data.

 

2. To make statements regarding

probabilities as to the occurrence

of a particular score, as in games

of chance.

 

3. To make statements regarding

the reliability of a single mean

 

4. To make statements regarding

the reliability of the difference

between two means.

 

 

 

Understanding the concepts in

number 4 above is the basis for

understanding the concepts of all

statistical tests. Also, as we said

before, there is a continuity in the

process of our understanding of

statistics. 

 

It is like a fairy tale. You must

know the full story, starting from

the beginning and step by step

reach the end, in order to make

sense. So keep alert!

 

Story 7 — Variance and standard deviation

Variance and standard deviation - intuitive level

Variance and standard deviation - formal level

Drama

Mathematical sweat

 

Susan Bolles is a new professor of

Psychology at Goatshead College.

Her chairman, sorry, chairperson, Dr.

Alexa Terrorvski, assigned her the

introductory psychology class of

1000 students. The first midterm

exam has just taken place. There

were 100 questions, 1 point each.

The exam papers were

computer-graded, all 1000 of them.

Dr. Terrorvski wants to know how

the class did, so she asks Susan.

Susan says that the mean (the

average) was 60.

 

Dr. Terrorvski wants to know more,

how many people scored close to

100, and how many people scored

close to zero. Susan walks up to the

pile of exam sheets and starts

reading the scores: 48, 30, 70, 99,

53 …… Papers are spread out from

the middle point, the average.

Realizing that this would take a good part

of the day, Dr. Terrorvski shouts out:

 

There must be a better way!

I will tell you what. Take this pile out to the

stadium, put it down in the middle of the

stadium. Mark this point 60 (your average

score). Take a step and mark this point 61.

Another step, mark this 62, all the way to

100. Return to mark 60 and take a step in

the opposite direction. Mark this point

59.Take another step, and mark this 58,

repeat all the way to mark 0. Now return to

your pile of exam sheets and pick up an

exam sheet. Read the score and walk to the

point it corresponds to on the markings you

made. I will return in an hour to see how

the exam went.

 

When Dr. Terrorvski returns she finds

Susan drenched in sweat and panting

vigorously. There is a long line of white

sheets of paper on both sides of the

point that marks 60, the average.

Susan picks up another exam sheet

which had the score of 3 and begins

walking. 59. 58. 57….

 

Enough! Terrorvski shouts. This is a

mess. Look at how many papers are

spread out away from the mean,

students have scored low scores, all

the way down to zero. So many

scores, so many students are very far

from 60, the mean. There is a big

distance of many scores from the

mean. You need to be more effective

in teaching your students, even the

weak ones, Dr. Terrorvski said, and

marched out of the stadium.

 

Susan went to her office and tried

to get a better picture of the

situation. Rather than walking

away from the point of the mean,

she calculated the distance of

each score from the mean. Score

40. Deviation from the mean: -20.

Score 65. Deviation from the mean:

+5. And so on. At the end she

added the absolute values of these distances and 

she found the total absolute distance.

That was the distance she had to

walk in the stadium!

 

In the second midterm, the mean

was again 60. This time she did

not go out in the stadium. She

simply found out the distance of

each score from the mean. She

added the absolute values of these distances

and was pleased to see that the

total absolute distance was very small.

Surely, Dr. Terrorvski would not

yell at her this time. Students

scored close to the mean, there

were very few low scores. The

scores this time were not spread

out all over the place away from

the mean.

 

The symbol for variance is s2.

The formula for variance is

 

 

\[ s^2=\frac{\sum (X-\bar{X})^2}{n-1} \]

 

 

We read this as follows:

Variance equals the sum of

squared deviations of each score

from the mean, divided by how

many scores went into the

calculation.

 

Why squared? Why square the

deviation, you say.

 

The sum of deviations from the

mean always, in all cases, equals

0. That is why we square each

deviation to prevent this. You

should know that in all sciences,

for the purpose of meaningful

analysis, we may transform our

data by squaring them, or

expressing them as logarithms,

and so on. This does not change

the relation of scores amongst

themselves. 

 

Why divide by n-1 for a sample?

 

You understand that if in one case

we have large scores, and in

another small scores, the sum of

the deviations from the mean will

be large in the first instance, and

small in the second instance. If we

want to compare the spread of the

scores in the two instances, we

have to average each of these

sums of deviations. I hope you

understand this.

 

For example, in order to compare

the income of New Yorkers to

that of Chicagoans, we must

average the total income of New

Yorkers and Chicagoans.

The numerator of this formula

 

 

\[ \sum (X-\bar{X})^2 \]

 

 

is the sum of the difference of

each score squared, or raised to

the second power. More formally,

we say the numerator of the
variance formula is the sum of

squared deviations of each score from the

mean, squared. In statistical jargon

 we say: 

 

Sum of squares, or SS.

 

The numerator of the variance

formula is the

Sum of Squares,

or SS

 

For a sample, the denominator is n-1; for

an entire population, the denominator is N.

This converts the sum of squared deviations

into the appropriate sample or population

variance.

 

So, the variance formula is the

average of the sum of squared

deviations, or, in statistical jargon,

the mean squares or MS, for

short. 

 

 

Variance is also called

mean squares

or

MS

 

 

 What is standard deviation? you say.

 

To calculate the standard deviation

we take the square root of

variance. Simple.

 

You do not need to know how to

calculate the square root of a

number. Not in the age of

computers. Anyone can learn to

do simple arithmetic. The

challenge is to understand

concepts of statistical and

mathematical operations.

 

The formula for standard deviation

is:

 

 

\[ s=\sqrt{s^2} \]

 

 

Do not worry about the -1 in the

denominator. The n changes

depending on whether we deal

with samples or an entire

population. Remember our goal 

here is to understand the concepts

of statistics and want to avoid

getting stuck in compulsive

swamps.

 

We use n-1 when we work with samples.

We N without -1 when we refer to population.

 

Now that we have removed the

mystery of standard deviation of

the normal distribution, we return

to it.

 

Remember this is not just a curve,

it is Goddess Normal Curve. Glory

to NC in the highest!

 

Appendix 11 — The F-table

Appendix 11: The F-Table (Critical Values at α = 0.05)

This table gives critical F-values for the F-distribution at the 5% significance level (one-tailed), used in ANOVA to test if group means differ significantly.

How to Use the F-Table

  • Left column: Degrees of freedom for the denominator (df₂ = df_within or error term).
  • Top row: Degrees of freedom for the numerator (df₁ = df_between or treatment term).
  • Find intersection value → that's the critical F.
  • If your calculated F > critical value → reject null hypothesis (significant difference at p < 0.05).

F Critical Values Table (α = 0.05)

F Critical Values (α = 0.05)
df₂ \ df₁12345678910
1161.45199.50215.71224.58230.16233.99236.77238.88240.54241.88
218.5119.0019.1619.2519.3019.3319.3519.3719.3819.40
310.139.559.289.129.018.948.898.858.818.79
47.716.946.596.396.266.166.096.046.005.96
56.615.795.415.195.054.954.884.824.774.74
65.995.144.764.534.394.284.214.154.104.06
75.594.744.354.123.973.873.793.733.683.64
85.324.464.073.843.693.583.503.443.393.35
95.124.263.863.633.483.373.293.233.183.14
104.964.103.713.483.333.223.143.073.022.98
114.843.983.593.363.203.093.012.952.902.85
124.753.893.493.263.112.992.912.852.802.75
134.673.813.413.183.032.922.832.772.712.67
144.603.743.343.112.962.852.762.702.652.60
154.543.683.293.062.902.792.712.642.592.54
164.493.633.243.012.852.742.662.592.542.49
174.453.593.202.962.812.702.612.552.492.45
184.413.553.162.932.772.662.582.512.462.41
194.383.523.132.902.742.632.542.482.422.38
204.353.493.102.872.712.602.512.452.392.35
214.323.473.072.842.682.572.492.422.372.32
224.303.443.052.822.662.552.462.402.342.30
234.283.423.032.802.642.532.442.372.322.27
244.263.403.012.782.622.512.422.362.302.25
254.243.392.992.762.602.492.402.342.282.24
264.233.372.982.742.592.472.392.322.272.22
274.213.352.962.732.572.462.372.312.252.20
284.203.342.952.712.562.452.362.292.242.19
294.183.332.932.702.552.432.352.282.222.18
304.173.322.922.692.532.422.332.272.212.16
404.083.232.842.612.452.342.262.202.152.11

Example: One-Way ANOVA

Experiment: 3 groups, 10 subjects each → df_between = 2, df_within = 27. Calculated F = 12.54.

From table: df₂ = 27, df₁ = 2 → critical F ≈ 3.35.

Since 12.54 > 3.35 → significant difference (p < 0.05). Reject null hypothesis: At least one group mean differs.

Tip: For more precision, different α levels (e.g., 0.01), or larger df, use statistical software (Excel: F.INV.RT, Google Sheets, R: qf()). See Appendix 5 for technology tips.

Related: Lesson 7 — Analysis of Variance (ANOVA)Part 7 — The Storyteller Statistician

Appendix 10 — The t-table

Appendix 10: The t-Table (Critical Values for Student's t Distribution)

This appendix provides critical t-values for hypothesis testing (e.g., one-sample or independent t-tests) at common significance levels. Use it to determine if your calculated t-statistic indicates a significant difference (reject null hypothesis when |t| > critical value).

How to Use the t-Table

  • Left column: Degrees of freedom (df) — usually n-1 for one-sample or n₁ + n₂ - 2 for two independent samples.
  • Top row: Significance level (p, one-tailed or two-tailed depending on test).
  • Find the intersection value = critical t.
  • If your |calculated t| > critical t → reject null hypothesis (significant at that p level).
  • For two-tailed tests, use p/2 values (e.g., for α=0.05 two-tailed, use 0.025 column).

Example 1 (from page content)

Two independent groups, 10 subjects each → df = 18.

Calculated t = 4.51.

For a one-tailed test at α = 0.05, df = 18 gives critical t = 1.734. For a two-tailed test at α = 0.05, use the 0.025 column, giving critical t = 2.101.

In either case, 4.51 exceeds the critical value, so the result is significant (p < 0.05). Reject null hypothesis: means differ.

Example 2

One-sample t-test, n = 25 → df = 24.

Calculated t = 2.15.

Look up df = 24, p = 0.05 → critical t = 1.711 (one-tailed) or use 0.025 column = 2.064 (two-tailed).

If one-tailed: 2.15 > 1.711 → significant. If two-tailed: 2.15 > 2.064 → significant.

t Critical Values Table

Student's t Critical Values
df0.100.050.0250.010.0050.001
13.0786.31412.70631.82163.657318.31
21.8862.9204.3036.9659.92522.326
31.6382.3533.1824.5415.84110.215
41.5332.1322.7763.7474.6047.173
51.4762.0152.5713.3654.0325.893
61.4401.9432.4473.1433.7075.208
71.4151.8952.3652.9983.4994.782
81.3971.8602.3062.8963.3554.499
91.3831.8332.2622.8213.2504.296
101.3721.8122.2282.7643.1694.143
111.3631.7962.2012.7183.1064.024
121.3561.7822.1792.6813.0553.929
131.3501.7712.1602.6503.0123.852
141.3451.7612.1452.6242.9773.787
151.3411.7532.1312.6022.9473.733
161.3371.7462.1202.5832.9213.686
171.3331.7402.1102.5672.8983.646
181.3301.7342.1012.5522.8783.610
191.3281.7292.0932.5392.8613.579
201.3251.7252.0862.5282.8453.552
211.3231.7212.0802.5182.8313.527
221.3211.7172.0742.5082.8193.505
231.3191.7142.0692.5002.8073.485
241.3181.7112.0642.4922.7973.467
251.3161.7082.0602.4852.7873.450
261.3151.7062.0562.4792.7793.435
271.3141.7032.0522.4732.7713.421
281.3131.7012.0482.4672.7633.408
291.3111.6992.0452.4622.7563.396
301.3101.6972.0422.4572.7503.385
401.3031.6842.0212.4232.7043.307
601.2961.6712.0002.3902.6603.232
1.2821.6451.9602.3262.5763.090

Tip: For exact p-values or larger df, use software (Excel: T.INV.2T, Google Sheets, R: qt()). See Appendix 5 for technology tips.

Related: Lesson 6 — The t-testLesson 7 — Analysis of Variance (ANOVA)

Appendix 9 — The normal distribution table

Appendix — The Normal Curve Table (Z-Table)

The Z-table (Standard Normal Distribution table) gives the area under the curve from the mean (z = 0) to a given z-score; it is not a left-tail cumulative-probability table. It is used in high school statistics to find probabilities, confidence intervals, and critical values for normal distribution problems.

How to Use the Z-Table

  • Left column: The z-score integer and first decimal (e.g., 1.9).
  • Top row: The second decimal place (0.00 to 0.09).
  • Find intersection → area from mean to that z-score (proportion of the distribution).
  • For negative z-scores, use symmetry (area is the same as positive).
  • For probability beyond z (tail), subtract from 0.5 (or 1 for two-tailed).

Example 1 (from page content)

z = 1.90 → row 1.9, column 0.00 → 0.4713. For z = 1.96, use row 1.9, column 0.06 → 0.4750.

Thus 47.13% lies between the mean and z = 1.90, and 47.50% lies between the mean and z = 1.96.

Example 2

Find probability that a score is less than z = 1.28 (e.g., for 90th percentile).

Row 1.2, column 0.08 → 0.3997.

Area from mean to z = 1.28 is 0.3997 → total area below z = 0.5 + 0.3997 = 0.8997 (≈90%).

Standard Normal (Z) Table — Cumulative Probabilities

Cumulative Probabilities from Mean to z (Standard Normal Distribution)
z0.000.010.020.030.040.050.060.070.080.09
0.00.00000.00400.00800.01200.01600.01990.02390.02790.03190.0359
0.10.03980.04380.04780.05170.05570.05960.06360.06750.07140.0753
0.20.07930.08320.08710.09100.09480.09870.10260.10640.11030.1141
0.30.11790.12170.12550.12930.13310.13680.14060.14430.14800.1517
0.40.15540.15910.16280.16640.17000.17360.17720.18080.18440.1879
0.50.19150.19500.19850.20190.20540.20880.21230.21570.21900.2224
0.60.22570.22910.23240.23570.23890.24220.24540.24860.25170.2549
0.70.25800.26110.26420.26730.27040.27340.27640.27940.28230.2852
0.80.28810.29100.29390.29670.29950.30230.30510.30780.31060.3133
0.90.31590.31860.32120.32380.32640.32890.33150.33400.33650.3389
1.00.34130.34380.34610.34850.35080.35310.35540.35770.35990.3621
1.10.36430.36650.36860.37080.37290.37490.37700.37900.38100.3830
1.20.38490.38690.38880.39070.39250.39440.39620.39800.39970.4015
1.30.40320.40490.40660.40820.40990.41150.41310.41470.41620.4177
1.40.41920.42070.42220.42360.42510.42650.42790.42920.43060.4319
1.50.43320.43450.43570.43700.43820.43940.44060.44180.44290.4441
1.60.44520.44630.44740.44840.44950.45050.45150.45250.45350.4545
1.70.45540.45640.45730.45820.45910.45990.46080.46160.46250.4633
1.80.46410.46490.46560.46640.46710.46780.46860.46930.46990.4706
1.90.47130.47190.47260.47320.47380.47440.47500.47560.47610.4767
2.00.47720.47780.47830.47880.47930.47980.48030.48080.48120.4817
2.10.48210.48260.48300.48340.48380.48420.48460.48500.48540.4857
2.20.48610.48640.48680.48710.48750.48780.48810.48840.48870.4890
2.30.48930.48960.48980.49010.49040.49060.49090.49110.49130.4916
2.40.49180.49200.49220.49250.49270.49290.49310.49320.49340.4936
2.50.49380.49400.49410.49430.49450.49460.49480.49490.49510.4952
2.60.49530.49550.49560.49570.49590.49600.49610.49620.49630.4964
2.70.49650.49660.49670.49680.49690.49700.49710.49720.49730.4974
2.80.49740.49750.49760.49770.49770.49780.49790.49790.49800.4981
2.90.49810.49820.49820.49830.49840.49840.49850.49850.49860.4986
3.00.49870.49870.49870.49880.49880.49890.49890.49890.49900.4990

Tip: For negative z-scores, the area is the same (symmetry). For tail probabilities, subtract from 0.5 (one-tailed) or 1 (two-tailed). Use software for exact values (Excel: NORM.S.DIST, Google Sheets, R: pnorm()). See Appendix 5 for technology tips.

Related: Lesson 4 — The Standard Normal CurvePart 7 — The Storyteller Statistician

About | High School Statistics (Pre-College)

About This Textbook

Statistics for High School Students: Pre-College is a free, broad-ranging, and interactive online textbook written by Dr. Michael Nikoletseas—a professor and researcher with numerous publications in neuroscience, philosophy of science, and mathematics.  Using mainly elementary arithmetic, straightforward formulas, and plain English, this resource is designed to be highly accessible. Despite its simplicity, it covers both elementary and advanced statistics topics, as well as modern data science concepts.

Mission & Vision

Our mission is to deliver a statistics textbook that:

  • supports students across a wide range of disciplines (from social and behavioral sciences to engineering and mathematics) to acquire a deep understanding of statistical reasoning, not just procedural techniques;
  • presents key statistical concepts in a manner that bridges theory and practice, emphasizing interpretive insight (“what does this mean?”) alongside computational method;
  • adopts an open mindset toward pedagogy: the site is structured for readability, modular use (individual chapters may be used independently if desired), and easy updates as the field evolves;
  • integrates modern elements—resampling, simulation, machine learning prelude, robust inference—while preserving the classical foundations (distributions, hypothesis testing, ANOVA, regression) so students are well‐grounded for further work.

Who This is For

This textbook is ideal for:

  • high school students preparing for biology or social science majors.
  • students in a one- or two-semester introductory statistics sequence who want more than formula memorization;
  • non‐mathematics majors (e.g. philosophy of science) who need to understand how to interpret and apply statistical reasoning in their discipline;
  • mathematics or statistics majors seeking a readable, web‐enabled resource that complements more formal references;
  • educators who want a ready‐to-use, modular, up‐to‐date resource for their course, including figures, examples, and modern topics.

Author & Credentials

Dr Michael Nikoletseas is the author of this textbook and brings a unique interdisciplinary background: his published works span neuroscience, philosophy of science, and mathematics, and are held in leading academic libraries (Harvard, Oxford, Princeton). His ambition with this text is to raise the bar for clarity, coherence, and depth in undergraduate statistics education.

With this online text, he applies the same analytical rigor he uses in his philosophical and mathematical writing: clear definitions, structured exposition, precise notation, and an emphasis on the limits of inference and interpretation (a theme that resonates with his broader work in epistemology).

Structure of the Textbook

The book is arranged into chapters each designed to stand on its own while also fitting into an integrated whole. Typical chapters will proceed in this order:

  1. Introduction & motivation
  2. Essential theory and notation for mathematics and formulas)
  3. Detailed examples and figures (copyable images for instructor use)
  4. Worked problems, with step-by-step solutions and commentary
  5. Live self-test quizzes
  6. Ask questions in each chapter
  7. Advanced topics, extensions, and links (for students preparing for further study)

Current chapters already include: descriptive statistics, probability, distributions, the normal distribution, hypothesis testing, t‐tests, one‐way and multi‐way ANOVA (including mixed designs and post-hoc comparisons), resampling and simulation, machine learning foundations, and big data computational statistics.

Contact & Feedback

Your feedback is valuable. Should you spot an error, have a suggestion for improvement or want to request supplementary material use Feedback on main menu. Use the contact form below each chapter to ask questions. 

Acknowledgements

The creation of this textbook has drawn on countless influences—from classical mathematics and modern statistics pedagogy to insights from neuroscience, philosophy of science, and epistemology. Special thanks to readers and educators who engage with the text, write with questions, and propose improvements. Together we advance statistical literacy and interpretive clarity.

Thank you for visiting StatisticsTextbook.com. May this textbook serve you well in your statistical journey.

— Michael Nikoletseas

Students

For Students: How to Use statisticstextbook.com

A simple guide for starting, studying in order, and reviewing.

Audience: Pre-college and high school students

1. What This Site Is

statisticstextbook.com is a free, page-by-page statistics textbook. You can read it in order like a print book, or use it as a reference when you need help with a topic.

Most students do best by moving from the foundations (data, variability, probability) into core tests (t-tests and ANOVA), and then into modern topics (resampling, big data, and an introduction to machine learning).

2. How to Use This Textbook

  1. Start with the first lesson.
  2. Follow the Next / Previous links. Each lesson ends with navigation links so you can keep the correct order without guessing what comes next.
  3. Keep a small “definitions” page in your notes. Write down the meaning of key terms (mean, variance, standard deviation, probability, distribution) as you encounter them.
  4. For each test, practice three skills. (1) what the question is, (2) the computation, (3) the interpretation in words.
  5. Use the review pages when you get stuck.

3. Reading the Math

Formulas are displayed with MathJax so they stay clear on different screens. If a formula looks unfamiliar, read it slowly and connect each symbol to a meaning in words.

4. Why This Format Helps

  • Clear sequence: lessons build from basic ideas to core tests.
  • Readable math: formulas render cleanly across devices.
  • Study-friendly: minimal distractions and no sign-in required.
  • Open access: free to use for learning and review.

5. Summary

Use the textbook in order if you are learning statistics for the first time, and use it as a reference when you need a quick explanation or a worked example. If you study steadily and keep your own notes of definitions and interpretations, the material becomes much easier over time.

© 2025. This page uses MathJax with LaTeX delimiters \(…\) and \[…\] in Drupal Full HTML.

Mixed (Split-Plot) ANOVA

mixed anova layout
mixed anova mean profile
partitioning variance
f distribution
split-plot interaction

Goal. Test a between-subjects factor (Group: Drug vs. Placebo) and a within-subjects factor (Time: Weeks 1–3), plus their interaction, on exam scores.

Design & Experiment

  • Between-subjects factor: Group = {Drug, Placebo}
  • Within-subjects factor: Time = {Week 1, Week 2, Week 3}
  • Balanced: 8 participants per group (\(s_g=8\)), 3 repeated measures per participant (\(k=3\)).

Participants are randomly assigned to Drug or Placebo. The same exam is given at Week 1, Week 2, and Week 3.

Figure 1: Mixed design layout (Drug vs Placebo × Weeks 1–3).


Data

Group: Drug (8 participants × 3 weeks)

SubjectW1W2W3Row sumRow mean
D170747822274.00
D269737721973.00
D371757922575.00
D472768022876.00
D568727621672.00
D670747822274.00
D773778123177.00
D871768022775.67
Column sums564597629Group sum = 1790Group mean \( \bar X_{\text{Drug}} = 1790/24 = 74.5833 \)

Group: Placebo (8 participants × 3 weeks)

SubjectW1W2W3Row sumRow mean
P170717221371.00
P269707121070.00
P371727321672.00
P472737421973.00
P568697020769.00
P670717221371.00
P769707121070.00
P871727321672.00
Column sums560568576Group sum = 1704Group mean \( \bar X_{\text{Plac}} = 1704/24 = 71.0000 \)

Totals. Grand sum = 1790 + 1704 = 3494, total observations \(N = 16\times3 = 48\), grand mean \( \bar X = 3494/48 = 72.7917\).

Figure 2: Mean profiles over weeks (Drug rises sharply; Placebo ~ flat).


Step 1 — Marginal Means

By Time (across both groups; 16 participants each week): \[ \bar X_{\text{W1}}=\tfrac{1124}{16}=70.2500,\qquad \bar X_{\text{W2}}=\tfrac{1165}{16}=72.8125,\qquad \bar X_{\text{W3}}=\tfrac{1205}{16}=75.3125, \] where column sums are \(1124, 1165, 1205\).

By Group (across all weeks): \[ \bar X_{\text{Drug}}=74.5833,\qquad \bar X_{\text{Placebo}}=71.0000. \]


Step 2 — Sums of Squares (SS)

Decompose total variability into Between-Subjects and Within-Subjects parts.

2A. Total

\[ SS_{\text{total}}=\sum (X_{igt}-\bar X)^2=\mathbf{527.9167}. \]

2B. Between-Subjects

Let each subject’s mean be \(\bar X_{i\cdot}\). Then \[ SS_{\text{BS-total}}=k\sum_{i=1}^{16}(\bar X_{i\cdot}-\bar X)^2=\mathbf{247.2500}. \] Split into Group and Subjects-within-Group: \[ SS_{\text{Group}}=k\sum_{g} n_g(\bar X_{g\cdot\cdot}-\bar X)^2=\mathbf{154.0833}, \] \[ SS_{\text{Subj}(g)}=k\sum_{i\in g}(\bar X_{i\cdot}-\bar X_{g\cdot\cdot})^2=\mathbf{93.1667}. \]

2C. Within-Subjects

\(SS_{\text{WS-total}}=SS_{\text{total}}-SS_{\text{BS-total}}=\mathbf{280.6667}.\)

Decompose into Time, Group×Time, and residual Error: \[ SS_{\text{Time}}=s\sum_{t}(\bar X_{\cdot\cdot t}-\bar X)^2=\mathbf{205.0417}, \] \[ SS_{\text{Group}\times\text{Time}} =\sum_{g,t} n_g\Big(\bar X_{g\cdot t}-\bar X_{g\cdot\cdot}-\bar X_{\cdot\cdot t}+\bar X\Big)^2 =\mathbf{75.0417}, \] \[ SS_{\text{Error(WS)}}=SS_{\text{WS-total}}-SS_{\text{Time}}-SS_{\text{G}\times\text{T}} =\mathbf{0.5833}. \]

Figure 3: Partitioning diagram (Between: Group + Subj(Group); Within: Time + G×T + Error).


Step 3 — Degrees of Freedom (df) & Mean Squares (MS)

\[ \begin{aligned} &df_{\text{Group}}=g-1=1,\qquad df_{\text{Subj}(g)}=N_s-g=16-2=14,\\ &df_{\text{Time}}=k-1=2,\qquad df_{\text{G}\times\text{T}}=(g-1)(k-1)=2,\\ &df_{\text{Error(WS)}}=(N_s-g)(k-1)=(16-2)\times2=28,\\ &df_{\text{Total}}=Nk-1=48-1=47. \end{aligned} \]

\[ \begin{aligned} &MS_{\text{Group}}=\frac{SS_{\text{Group}}}{df_{\text{Group}}}= \frac{154.0833}{1}= \mathbf{154.0833},\qquad MS_{\text{Subj}(g)}=\frac{93.1667}{14}= \mathbf{6.6548},\\ &MS_{\text{Time}}=\frac{205.0417}{2}= \mathbf{102.5208},\qquad MS_{\text{G}\times\text{T}}=\frac{75.0417}{2}= \mathbf{37.5208},\\ &MS_{\text{Error(WS)}}=\frac{0.5833}{28}= \mathbf{0.02083}. \end{aligned} \]


Step 4 — F Tests & p-values

Between-subjects test: \[ F_{\text{Group}}=\frac{MS_{\text{Group}}}{MS_{\text{Subj}(g)}}=\frac{154.0833}{6.6548}= \mathbf{23.1538}, \quad df=(1,14),\quad p\approx \mathbf{0.00028}. \]

Within-subjects tests: \[ F_{\text{Time}}=\frac{MS_{\text{Time}}}{MS_{\text{Error(WS)}}} =\frac{102.5208}{0.02083}= \mathbf{4921.0},\quad df=(2,28),\quad p\ll 10^{-20}. \] \[ F_{\text{G}\times\text{T}}=\frac{MS_{\text{G}\times\text{T}}}{MS_{\text{Error(WS)}}} =\frac{37.5208}{0.02083}= \mathbf{1801.0},\quad df=(2,28),\quad p\ll 10^{-20}. \]

Figure 4: F distributions with observed statistics marked.


Mixed ANOVA Summary Table

SourceSSdfMSFp
Between: Group154.08331154.083323.15380.00028
Between: Subjects within Group93.1667146.6548
Within: Time205.04172102.52084921.0< 1e-20
Within: Group × Time75.0417237.52081801.0< 1e-20
Within: Error (Subj×Time within Group)0.5833280.02083
Total527.916747

Interpretation

Group: Drug > Placebo overall (significant between-subjects effect).
Time: Scores increase across weeks (strong within-subjects effect).
Group × Time: The Drug group improves sharply week-to-week while the Placebo group changes little (significant interaction).

Figure 5: Interaction plot showing non-parallel lines (Drug rising; Placebo flat).

Assumptions (checklist)

  • Independence between subjects; correct grouping.
  • Approximate normality within each Group×Time cell.
  • Homogeneity of variance across groups (between-subjects).
  • Sphericity for the within-subject factor Time (apply Greenhouse–Geisser/Huynh–Feldt corrections if violated).

Note: The residual within-subject error is intentionally small in this teaching dataset, so the Time and G×T F values are very large. Real data typically have larger residual variability.

Practice self-test quiz

In the space below, please find practice problems and self-test quizzes. For full access, please signup free.

Repeated-Measures ANOVA

rm profile
rm sem
rm partitioning var
f distrib
rm sphericity

Goal. Test whether performance changes across four conditions measured on the same participants.

Design & Experiment

  • Within-subjects factor: Condition with 4 levels (C1, C2, C3, C4).
  • s = 8 participants measured in k = 4 conditions ⇒ total observations \(N = s \times k = 32\).
  • Example context: the same students take four weekly quizzes after different study activities.

Figure 1: Profile plot (each subject as a line across the four conditions).


Data

Scores (rows = participants S1–S8; columns = conditions C1–C4):

SubjectC1C2C3C4Row sumRow mean
S17074758130075.00
S27375788230877.00
S36873737829273.00
S47479818531979.75
S57174788230576.25
S67072767829674.00
S77377808431478.50
S87477808431578.75
Column sums573601621654Grand sum = 2449Grand mean \( \bar X = 2449/32 = 76.53125 \)

Figure 2: Means ± SEM for C1–C4 (bar/line).


Step 1 — Condition Means (and sample variances)

\[ \begin{aligned} \bar X_{\mathrm{C1}} &= 573/8 = 71.625, \quad & s^2_{\mathrm{C1}} &= 4.8393 \\ \bar X_{\mathrm{C2}} &= 601/8 = 75.125, \quad & s^2_{\mathrm{C2}} &= 5.5536 \\ \bar X_{\mathrm{C3}} &= 621/8 = 77.625, \quad & s^2_{\mathrm{C3}} &= 7.6964 \\ \bar X_{\mathrm{C4}} &= 654/8 = 81.750, \quad & s^2_{\mathrm{C4}} &= 7.0714 \end{aligned} \]


Step 2 — Sums of Squares

Notation: \(s=8\) subjects, \(k=4\) conditions, grand mean \( \bar X = 76.53125\).

2A. Total

\[ SS_{\text{total}}=\sum_{i=1}^{s}\sum_{j=1}^{k}\bigl(X_{ij}-\bar X\bigr)^2 =\mathbf{611.96875}. \]

2B. Conditions (Treatment)

\[ SS_{\text{cond}}= s \sum_{j=1}^{k}\bigl(\bar X_{\cdot j}-\bar X\bigr)^2 = 8 \left[(71.625-76.53125)^2 + (75.125-76.53125)^2 + (77.625-76.53125)^2 + (81.75-76.53125)^2\right] =\mathbf{435.84375}. \]

2C. Subjects

\[ SS_{\text{subj}}= k \sum_{i=1}^{s}\bigl(\bar X_{i\cdot}-\bar X\bigr)^2 = 4 \sum_{i=1}^{8}\bigl(\bar X_{i\cdot}-76.53125\bigr)^2 =\mathbf{162.71875}. \]

2D. Error (Residual)

\[ SS_{\text{error}}= SS_{\text{total}} - SS_{\text{cond}} - SS_{\text{subj}} = 611.96875 - 435.84375 - 162.71875 =\mathbf{13.40625}. \]

Figure 3: Partitioning variance diagram (Total → Conditions + Subjects + Error).


Step 3 — Degrees of Freedom & Mean Squares

\[ \begin{aligned} df_{\text{cond}} &= k-1 = 3, \\ df_{\text{subj}} &= s-1 = 7, \\ df_{\text{error}} &= (s-1)(k-1) = 7\times3 = 21, \\ df_{\text{total}} &= sk-1 = 31. \end{aligned} \]

\[ MS_{\text{cond}} = \frac{SS_{\text{cond}}}{df_{\text{cond}}} =\frac{435.84375}{3}=\mathbf{145.28125},\qquad MS_{\text{error}} = \frac{SS_{\text{error}}}{df_{\text{error}}} =\frac{13.40625}{21}=\mathbf{0.6383928571}. \]


Step 4 — Test Statistic & p-value

\[ F = \frac{MS_{\text{cond}}}{MS_{\text{error}}} = \frac{145.28125}{0.6383928571} =\mathbf{227.5734}. \] With \(df_1=3\) and \(df_2=21\), this is extremely large. The right-tail p-value is effectively \(p \lt 10^{-12}\) (i.e., \(p \ll .001\)).

Figure 4: F distribution with observed F marked and right-tail region shaded.


Repeated-Measures ANOVA Summary Table

SourceSSdfMSFp
Conditions (within)435.843753145.28125227.5734< 1e-12
Subjects162.71875723.24554
Error (residual)13.40625210.63839
Total611.9687531

Interpretation

Mean performance increases steadily from C1 → C4, and the repeated-measures ANOVA shows a highly significant effect of Condition, \(F(3,21)=227.57,\, p\ll .001\). Follow-ups (e.g., paired t-tests with Bonferroni/Holm) can localize which pairs of conditions differ.

Assumptions (checklist)

  • Sphericity (equal variances of the differences between condition pairs). If violated, apply Greenhouse–Geisser or Huynh–Feldt correction to \(df\).
  • Approximately normal scores within each condition.
  • No carryover/fatigue effects that confound order (counterbalancing helps).

Figure 5: Sphericity concept sketch (pairwise difference variances).

Practice self-test quiz

In the space below, please find practice problems and self-test quizzes. For full access, please signup free.