The mean, median, and mode are all called "measures of central tendency" — they all describe where the centre of a dataset lies. But they tell different stories, and choosing the wrong one can seriously mislead your analysis.
The mean (average)
The mean is the sum of all values divided by the count. If five employees earn £30k, £35k, £38k, £42k, and £55k, the mean is (30 + 35 + 38 + 42 + 55) / 5 = £40k.
The mean is the most mathematically useful measure — it factors every value into the result, works with other statistical formulas, and provides the basis for standard deviation, regression, and most inferential statistics.
The problem: the mean is sensitive to outliers. Add a CEO earning £500k to that list of five employees and the mean jumps to £117k — a figure that doesn't represent most employees' experience at all. In skewed distributions, the mean is pulled toward the tail.
Use the mean when: your data is roughly symmetrically distributed and free of extreme outliers. Test scores, physical measurements, and many natural phenomena follow roughly normal distributions where the mean is the appropriate centre.
The median
The median is the middle value when data is sorted in order. If you have an odd number of values, it's the exact middle. If you have an even number, it's the average of the two middle values.
For the five employees above (£30k, £35k, £38k, £42k, £55k), the median is £38k. Add the £500k CEO and the median becomes the average of the 3rd and 4th values in the sorted list: (38 + 42) / 2 = £40k. The extreme value barely changed the median.
Use the median when: your data is skewed or contains outliers. House prices and income distributions are almost always reported as medians for this reason. A single billionaire in a small sample would make mean income misleading; the median is robust to that distortion.
The mode
The mode is the most frequently occurring value. A dataset can have no mode (if all values are unique), one mode (unimodal), or multiple modes (bimodal, trimodal, etc.).
For numerical data with many distinct values, the mode is rarely useful — the chance of any specific value occurring more than once is low. But for categorical data, the mode is often the most important statistic. "What is the most common product category purchased?" "What is the most frequently selected survey response?" These are mode questions.
Use the mode when: you're working with categorical data (product categories, survey responses, colours, countries), or when you specifically need to know the most common value in a discrete numerical dataset (size 10 shoes, 6-inch phone screens).
Which one should you report?
The safest approach is to report all three for any dataset you're seriously analysing, then choose the one that's most appropriate for your audience and question. The relationship between the three tells you something important:
- If mean ≈ median ≈ mode, the data is roughly symmetric. Report the mean.
- If the mean is significantly higher than the median, the data is right-skewed (a few very high values). Report the median.
- If the mean is significantly lower than the median, the data is left-skewed. Report the median.
- If you have categorical data, the mode is your primary measure.
A practical example: house prices
Suppose a neighbourhood has ten houses selling for: £180k, £190k, £195k, £200k, £205k, £210k, £215k, £220k, £240k, and £2,500k (a large estate).
- Mean: £435.5k — skewed heavily by the estate.
- Median: £207.5k — the value in the middle of the sorted list, much more representative of what most buyers will actually pay.
- Mode: none — all values are unique.
The median is clearly the right number to report here. Reporting the mean would give buyers a completely inaccurate picture of what to expect to pay.
Beyond central tendency
Whichever measure of centre you choose, always report it alongside a measure of spread — the standard deviation (for means), the interquartile range (for medians), or the full frequency distribution (for modes). Central tendency without spread is half the picture.