Passage
The authority of the average
Averages are often introduced as modest summaries. They do not claim to describe every case; they compress variation into a figure that can be compared. Yet once an average enters an institution, its modesty is easily forgotten.
It becomes a norm against which individuals are judged, a forecast around which resources are organised, or a threshold that determines who is considered exceptional. The calculation remains descriptive, while its administrative life becomes prescriptive. This transformation is clearest when the population being averaged is treated as stable.
A school may compare a pupil’s progress with an average derived from previous cohorts. The figure appears to provide context, but the cohorts were taught under different conditions and were themselves selected by attendance, assessment design and exclusion. The average does not merely summarise a group; it inherits the decisions that produced the group and carries them into a new judgement.
Critics sometimes respond by demanding more granular data. Instead of one national average, they propose averages for region, age, income or prior attainment. Such refinement can reveal inequalities hidden by aggregation, but it also creates smaller categories that appear more personally relevant.
The individual is now compared with a group that looks like them statistically, which can make the resulting expectation harder to question. Precision of resemblance is not the same as validity of inference. The authority of averages also depends on what happens to cases far from the centre.
In engineering, an average load may be useful only because design separately considers extremes. In health or transport policy, however, exceptional cases are often described as outliers and removed before the average is calculated. Sometimes this is methodologically justified; a broken sensor should not define normal traffic.
But a person whose journey or body does not fit the model is not a defective instrument. Exclusion may improve the stability of the statistic while worsening the institution’s ability to recognise whom it fails. There is a further complication.
People adjust their behaviour to the metrics built from averages. A call centre given an average handling-time target may shorten conversations whose complexity requires patience. A hospital monitored by average waiting time may move difficult cases between categories.
These responses are commonly dismissed as gaming, yet the metric has changed the environment it claims to measure. The average becomes accurate in a narrower sense because work is reorganised to satisfy its definition. None of this makes averaging inherently deceptive.
Collective decisions require compression; without it, every case remains incomparable and resources cannot be coordinated. The question is whether the summary preserves access to the variation it has compressed. A responsible average is accompanied by its distribution, the conditions of inclusion, the consequences of being distant from the centre and the reasons the category was created.
It should support interrogation rather than end it. This requires institutions to resist the rhetorical convenience of a single number. Decision-makers often ask for one figure because one figure can travel through a report, a headline and a budget meeting without carrying its uncertainties.
The statistic’s portability is part of its power. A more honest presentation may be less efficient: it may require several ranges, a description of missing cases and an admission that the relevant comparison changes with the decision being made. The aim is not to replace averages with individual stories, which carry their own selection effects and can be made falsely representative.
It is to keep the movement from summary to standard visible. An average should be treated as an answer to a specified question about a constructed group, not as the voice of a population speaking without mediation. The question of distance from the average is also asymmetric.
Being above a mean may be rewarded in one system and treated as suspicious in another; being below it can indicate need, failure or merely a different sequence of development. The figure does not contain these meanings. They arise from institutional purposes that are often imported into presentation through colour, ranking and language, then attributed back to the number as though its moral direction were mathematical.
Institutional dashboards intensify this effect by arranging several averages so that their visual alignment suggests commensurability, even when one summarises duration, another cost and a third a score whose scale was created for a different population. Once placed in adjacent columns, movement in each appears to describe a common performance trajectory, and the act of comparison recedes behind the neatness of the display. Responsible use must therefore disclose not only how each number was calculated but why these summaries are being asked to stand beside one another as evidence of the same thing.