How do I combine experimental probabilities from different sessions?

I’m practicing experimental probability with a red/blue spinner, and I’m stuck on how to report an overall probability when I ran the experiment in separate sessions.

Here’s what I did:
– Session 1: 20 spins, 9 red → 9/20 = 45%
– Session 2: 200 spins, 88 red → 88/200 = 44%
– Session 3: 500 spins, 190 red → 190/500 = 38%

Then I tried to get an overall experimental probability in two ways:
1) Average the three percentages: (45% + 44% + 38%) / 3 = 42.3%
2) Pool all results: total reds 9+88+190 = 287 out of 20+200+500 = 720 → 287/720 ≈ 39.9%

These don’t match, and I’m not sure which one is the “right” way or why. I feel like I’m mixing up how averages should work. My (possibly bad) analogy is: is this like averaging fuel economy over trips of different lengths (where you should weight by distance), or is it more like tasting three bowls of soup and just averaging the taste scores? I also wonder if I should be resetting the experimental probability each session, or if it’s fine to keep a running one.

One more detail: in Session 3 I stopped when I got tired, not after a fixed number I planned in advance. Does that kind of stopping rule affect the experimental probability I should report?

I thought experimental probability would get closer to a stable value as I do more spins, but my percentages went from 45% to 38%, which makes me doubt my method. Can someone walk me through step by step where my reasoning is going off and how I should properly combine batches?

2 Responses

  1. Love this question! Your fuel-economy analogy is spot on: to estimate one underlying probability from batches of different sizes, you should weight by how many spins each batch has. In probability-speak, the best combined estimate is the pooled proportion: total reds divided by total spins. Equivalently, it’s the weighted average of the session percentages with weights equal to the session sizes n. For your data that’s 287 reds out of 720 spins, 287/720 ≈ 39.9%. The unweighted average of 45%, 44%, and 38% treats a 20-spin session the same as a 500-spin session, which overemphasizes the noisier small batches.

    Here’s a tiny worked example to make the weighting idea pop. Suppose Session A has 1 spin with 1 red (100%), and Session B has 99 spins with 40 red (40%). The simple average is (100% + 40%)/2 = 70%, which wildly overstates red because it gives that lone spin equal voice. The pooled estimate uses all the data at once: (1 + 40) / (1 + 99) = 41/100 = 41%. That’s the same as the weighted average 100%×1/100 + 40%×99/100. In your case, the larger third session (38%) naturally pulls the overall downward; that’s expected and exactly what “more data has more say” should do. A great habit is to keep a running cumulative proportion (total reds so far / total spins so far); by the law of large numbers it will jitter less and settle as you collect more spins.

    About stopping rules: if you stopped a session because you got tired or ran out of time (i.e., the stopping point isn’t driven by the outcomes), your pooled proportion remains perfectly valid and unbiased. If instead you stop based on the outcomes themselves (for example, “I’ll stop as soon as I see 100 reds”), the naive proportion reds/spins can be slightly biased in small samples, though it’s still consistent and the effect fades with large counts. Bottom line: it’s fine not to “reset” each session-just keep pooling all the spins together for your experimental probability.

  2. Lovely setup! Think of all your spins as marbles in one big jar: if you want the overall experimental probability from everything you actually observed, you should pool the counts, i.e., total reds divided by total spins, which gives 287/720 ≈ 39.9%. Averaging the three percentages gives each session equal weight no matter how many spins happened there, which is like averaging fuel economy per trip without caring about trip length-cute, but it answers a different question: “what’s the average session’s red rate?” rather than “what’s the red rate per spin overall.” Both numbers can be reported, but the pooled one is the standard estimate of the underlying chance per spin. The big swing from 45% to 38% isn’t evidence of a mistake by itself-random wiggles calm down only when the total number of spins gets large, and Session 3 (with 500 spins) has more pull on the pooled estimate. About stopping: if you stopped when you got tired (not based on outcomes), that doesn’t bias the pooled proportion; the simple reds/total is still the natural estimate. If you had stopped based on outcomes (like “quit as soon as I see 50 reds”), the same proportion is still a reasonable point estimate, though the uncertainty behaves differently-tiny tangent, but the headline is you can still report reds/total. You don’t need to “reset” each session unless you suspect the spinner changed; keeping a running proportion is usually best, although averaging session percentages is fine if you specifically want to treat each session as one equal-sized snapshot. Out of curiosity, do you think the spinner could be drifting over time, or are you comfortable assuming it’s the same throughout?

Leave a Reply

Your email address will not be published. Required fields are marked *

Join Our Community

Ready to make maths more enjoyable, accessible, and fun? Join a friendly community where you can explore puzzles, ask questions, track your progress, and learn at your own pace.

By becoming a member, you unlock:

  • Access to all community puzzles
  • The Forum for asking and answering questions
  • Your personal dashboard with points & achievements
  • A supportive space built for every level of learner
  • New features and updates as the Hub grows