I’m trying to compute the Pearson correlation by hand for a small dataset and I’m clearly mixing things up. My data:
x = [1, 2, 3, 4, 5]
y = [2, 5, 7, 10, 12]
My (wrong) attempt: I took the average of the products, sum(xi*yi)/n. That gave me (2 + 10 + 21 + 40 + 60)/5 = 133/5 = 26.6, so I wrote down r = 26.6. Then I thought maybe correlation is the slope of the best-fit line; a quick fit gave me a slope around 2.5, so I tried calling that the correlation too. Both break the -1 to 1 rule, so I know I’m off.
Can someone point out the exact steps I should be doing for Pearson’s r here? Do I subtract the means first and then divide by standard deviations? Also, do I use n or n-1 in the denominators when I compute the standard deviations and the covariance?
Follow-up: If I multiply all y values by 10 or add a constant to y, should the correlation change? I assumed scaling would make it 10 times bigger because the products get bigger, but I’ve been told correlation is unitless. And if I reverse-code y by multiplying by -1, does r simply flip sign every time?
One last check: if the relationship is a bit curved (which this might be), is Pearson’s r still a reasonable summary, or should I be using something else?
















3 Responses
Nope-r isn’t the slope or the raw average of products; it’s the dot product of centered-and-scaled variables: r = Σ[(xi−x̄)(yi−ȳ)] / √(Σ(xi−x̄)^2 Σ(yi−ȳ)^2), and for your data x̄=3, ȳ=7.2, Σdx·dy=25, Σdx^2=10, Σdy^2=62.8, so r ≈ 25/√(10·62.8) ≈ 0.998.
Use n in the denominators (some folks use n−1 for SDs-hardly changes r), scaling or adding a constant to y won’t change r (multiplying by −1 just flips the sign), and if it’s curved try Spearman or a simple quadratic; refresher: https://www.khanacademy.org/math/statistics-probability/describing-relationships-quantitative-data/correlation-coefficient/a/calculating-correlation-coefficient-r
You’re not alone-I once did exactly what you did: I averaged the raw products, got some giant number, and wondered why my “correlation” was blasting past 1. The fix is to center first and then scale. Pearson’s r is r = sum[(xi − x̄)(yi − ȳ)] / sqrt{ sum[(xi − x̄)²] · sum[(yi − ȳ)²] }. For your data, x̄ = 3 and ȳ = 7.2. Then Sxy = 25, Sxx = 10, Syy = 62.8, so r = 25 / sqrt(10·62.8) ≈ 0.998-nice and high, but safely between −1 and 1. About n vs n−1: if you use n−1 in both the covariance and the standard deviations (the usual “sample” choice), or n in both (the “population” choice), it cancels out and r is the same either way-just be consistent. On your follow-up: adding a constant to y doesn’t change r at all; multiplying y by a positive constant leaves r unchanged; multiplying by a negative constant simply flips the sign of r. That’s the “unitless” part in action: scaling gets canceled out by the standard deviations. If the relationship is curved, Pearson’s r only measures the linear part, so it can be misleading; in that case consider Spearman’s rank correlation or fit a curve (e.g., a quadratic) and look at the residuals. When I finally saw this written as r = slope × (sd_x / sd_y), it all clicked for me-slope alone isn’t bounded, but once you scale by the spreads, you land in [−1, 1]. A nice walkthrough is here: https://www.khanacademy.org/math/statistics-probability/describing-relationships-quantitative-data/correlation-coefficient/v/pearson-correlation-coefficient
You’re super close-what you computed first was the average of the products E[XY], not the correlation. For Pearson’s r, center both variables first, then standardize the scales. Concretely: x̄ = 3 and ȳ = 7.2. Compute the cross-deviations sum∑(xi−x̄)(yi−ȳ) = 25, the x sum of squares ∑(xi−x̄)² = 10, and the y sum of squares ∑(yi−ȳ)² = 62.8. Then r = 25 / sqrt(10 · 62.8) ≈ 0.998, which fits the “between −1 and 1” rule nicely. The slope of the best-fit line isn’t the correlation; it’s b = ∑(xi−x̄)(yi−ȳ) / ∑(xi−x̄)² = 25/10 = 2.5, and it relates to r by b = r · (sy/sx). About n vs n−1: if you use n−1 for both the covariance and the standard deviations (the usual “sample” formulas), the factors cancel in r, so you’ll get the same answer as using n consistently. Scaling or shifting won’t change r: adding a constant to y or multiplying by a positive constant leaves r unchanged; multiplying by a negative constant just flips the sign of r. And yes, I also have to keep reminding myself correlation is unitless! If the relationship is curved, Pearson’s r only measures linear association; for monotonic but nonlinear trends, Spearman’s rank correlation is often better, and for clearly curved shapes you might fit a nonlinear model or transform variables. Hope this helps!