In a histogram, density describes how much data is concentrated within a particular range of values. It helps us understand how common values are in different parts of a distribution.
The easiest way to think about density in a histogram is:
The area of a bar represents the proportion of observations in that range.
For example, if a bar represents 20% of the area of the entire histogram, then about 20% of the observations fall within that range.
The height of the bar represents density. The height tells us how concentrated the observations are, taking the width of the bin into account.
The height of a histogram bar depends on both:
How much data is in the bin
How wide the bin is
This means that the height of a bar is not necessarily the proportion of observations in that bin.
Instead, we use the area of the bar to represent the proportion.
For a simple histogram where all bins have the same width, taller bars also contain more observations. But when bin widths are different, it is especially important to think about area rather than height.
Suppose a histogram has a bin representing ages from 20 to 30. The area of that bar is 0.25.
This means:
About 25% of the observations are between ages 20 and 30.
If another bin has an area of 0.10, then about 10% of the observations are in that range.
A density histogram is scaled so that the total area of all the bars is 1, representing 100% of the observations.
This makes it possible to compare distributions even when datasets have different numbers of observations.
Density helps us see where observations are concentrated within a distribution.
For example, a region with high density contains observations that are more concentrated than a region with low density.
When interpreting a histogram, look at the area of the bars to determine the proportion of observations in different ranges.
Density tells us how concentrated the data are, while the area of a histogram bar tells us the proportion of observations.
Remember:
Density → height
Proportion → area