Sound representation
Sound in the real world is a continuously varying pressure wave; a computer can only store a finite list of binary numbers. Digital sound bridges that gap by sampling: measuring the wave at regular moments and storing each measurement as an integer. This note covers the key terms (sampling, sampling rate, sampling resolution, analogue and digital data), how to calculate the size of a sound file, and how changing the sampling rate and resolution affects accuracy and file size, which Paper 1 tests in nearly every series.
Analogue and digital data
Analogue data is data that varies continuously: between any two values there are infinitely many others. A sound wave, a temperature or a voltage from a microphone are analogue.
Digital data is data that takes only discrete (separate) values, stored as binary numbers.
A microphone converts the pressure wave into an analogue electrical signal whose voltage rises and falls with the sound. An analogue-to-digital converter (ADC) then samples that voltage and outputs binary numbers. When the sound is played back, a digital-to-analogue converter (DAC) turns the stored numbers back into a varying voltage that drives a loudspeaker.
Sampling
Sampling is measuring the amplitude (height) of a sound wave at regular time intervals and recording each measurement as a binary number.
The sampling rate is the number of samples taken per second, measured in hertz (Hz).
The sampling resolution is the number of bits used to store each sample (also called the bit depth).
The sampling resolution decides how many different amplitude levels are available. With bits there are levels. Each measurement is rounded to the nearest available level, a process called quantisation. The difference between the true amplitude and the stored level is the quantisation error.
The graph shows a sound wave sampled once per time unit with a 3-bit sampling resolution, so there are levels, 0 to 7. Each dot is the stored, rounded value.
| Time | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| True amplitude | 3.30 | 4.99 | 6.10 | 6.22 | 5.33 | 3.72 | 1.97 | 0.69 | 0.31 | 0.98 | 2.46 |
| Stored level | 3 | 5 | 6 | 6 | 5 | 4 | 2 | 1 | 0 | 1 | 2 |
| Binary | 011 | 101 | 110 | 110 | 101 | 100 | 010 | 001 | 000 | 001 | 010 |
When this is played back, the output jumps between the stored levels. The digital version is a stepped approximation of the original: it misses what happens between samples, and each sample is slightly wrong because of rounding.
Improving the accuracy of a recording
There are two ways to make the digital version closer to the original wave.
Increase the sampling rate. Taking more samples per second means the wave is measured at more points, so faster changes (higher frequencies) are captured and the playback follows the original shape more closely.
Increase the sampling resolution. Using more bits per sample gives more amplitude levels, so each sample can be stored closer to its true value. Quantisation error is smaller and the range between the quietest and loudest sounds that can be represented (the dynamic range) is greater.
Both improvements cost space: the file size is directly proportional to each.
Not required by the syllabus, but it explains the numbers you see. To capture a frequency , you must sample at more than . Human hearing reaches about 20 kHz, so CD audio is sampled at 44.1 kHz. Telephone speech only needs frequencies up to about 4 kHz, so phone systems sample at 8 kHz.
Sound file size
Every second of sound contains (sampling rate) samples, each of (sampling resolution) bits, for each channel. Stereo has two channels (left and right), so it doubles the data.
Divide by 8 for bytes. As with images, this is an estimate: a real file also has a header (and may be compressed).
| Change | Effect on file size | Effect on sound |
|---|---|---|
| Double the sampling rate | Doubles | More accurate; higher frequencies captured |
| Increase resolution 8 → 16 bits | Doubles | More levels; less quantisation error; greater dynamic range |
| Mono → stereo | Doubles | Separate left and right channels |
| Halve the length | Halves | Shorter recording |
- Convert the sampling rate to Hz (44.1 kHz = 44 100 Hz) and the length to seconds.
- Multiply sampling rate × sampling resolution × time × channels to get bits.
- Divide by 8 to get bytes, then convert to the unit asked for.
- State the unit and note that the header is ignored.
A song lasts 3 minutes. It is recorded in stereo at a sampling rate of 44.1 kHz with a sampling resolution of 16 bits. Estimate the file size in MB.
Solution
Time: s. Sampling rate: Hz. Channels: 2.
Bits: .
Bytes: .
Size: (about ).
A voice memo is recorded in mono at 8 kHz with 8-bit samples. Calculate the size of a one-minute memo in kB.
Solution
bits bytes .
(8-bit samples are 1 byte each, so this is simply bytes.)
An audio CD holds 700 MB. CD audio uses 44.1 kHz, 16-bit stereo. Estimate the maximum playing time in minutes.
Solution
Bytes per second: bytes per second.
Seconds: s.
Minutes: minutes.
A podcast is recorded in mono at 16 kHz with 16-bit resolution. The producer considers changing to 48 kHz with 24-bit resolution. Calculate the factor by which the file size increases, and explain the effect on the recording.
Solution
File size is proportional to sampling rate × resolution.
Factor . The file becomes 4.5 times larger.
Effect: three times as many samples per second, so the wave is measured more often and higher frequencies are captured; 24 bits give levels rather than , so each sample is closer to the true amplitude (smaller quantisation error). The recording is a more accurate representation of the original sound.
Mixing up rate and resolution. Sampling rate is how often (samples per second). Sampling resolution is how precisely (bits per sample). Examiners penalise answers that swap them.
Forgetting stereo. If the question says stereo, multiply by 2.
Forgetting kHz. 44.1 kHz is 44 100 samples per second, not 44.1.
Saying a higher sampling rate makes sound "louder" or "clearer". Use precise language: a higher sampling rate gives a more accurate representation of the wave because it is measured more often; higher resolution reduces quantisation error.
Definitions are worth one mark each and examiners look for key words: sampling rate is "the number of samples taken per second"; sampling resolution is "the number of bits used to store each sample".
For "explain the impact of increasing the sampling resolution", give both the accuracy effect (more amplitude values available, each sample closer to the original, less quantisation error) and the file size effect (more bits per sample, larger file). Many candidates give only one side and lose a mark.
In calculations, write the full expression first, then simplify. No calculator is allowed: cancel factors of 8 early (16 bits = 2 bytes).
- Sound is analogue; it is stored digitally by sampling its amplitude at regular intervals.
- Sampling rate: samples per second (Hz). Sampling resolution: bits per sample.
- -bit resolution gives amplitude levels; rounding to a level is quantisation.
- Higher rate or resolution gives a more accurate recording and a bigger file.
- File size (bits) = rate × resolution × seconds × channels; divide by 8 for bytes.
- An ADC converts analogue to digital; a DAC converts back for playback.
Practice
- Explain the difference between analogue data and digital data.
- Define the terms sampling rate and sampling resolution.
- Calculate the number of amplitude levels available with a sampling resolution of 10 bits.
- A 10-second sound effect is recorded in mono at 22 050 Hz with 8-bit samples. Calculate the file size in bytes.
- A recording studio records at 48 kHz, 24-bit stereo. Calculate the size of a one-minute recording in MB.
- Describe two effects of reducing the sampling rate of a recording.
- A 5-minute interview is recorded in mono at 16 kHz with a 16-bit resolution. Calculate the file size in MiB, to one decimal place.
- A sound is recorded with a 2-bit sampling resolution. The true amplitudes at five sample points, on a scale where the levels are 0, 1, 2 and 3, are 0.4, 1.7, 2.9, 2.2 and 0.6. State the stored binary values and the largest quantisation error.
- A music streaming service offers "normal" quality (sampled at 22.05 kHz, 16-bit stereo) and "high" quality (44.1 kHz, 16-bit stereo), both uncompressed. Calculate the bit rate of each in kilobits per second, and explain why a user with a slow connection might choose normal quality.
- A student claims: "If I double the sampling rate and halve the sampling resolution, the file size stays the same, so the sound quality will be the same too." Evaluate this claim.
Answers
- Analogue data varies continuously and can take any value in a range; digital data takes only discrete values, stored as binary numbers.
- Sampling rate: the number of samples taken per second. Sampling resolution: the number of bits used to store each sample.
- levels.
- bytes.
- bytes .
- Fewer samples per second, so the wave is measured less often: the recording is a less accurate representation and high-frequency sounds may be lost. The file size is smaller (in proportion to the rate), so it needs less storage and bandwidth.
- bytes; .
- Rounded levels 0, 2, 3, 2, 1, stored as
00,10,11,10,01. Errors are 0.4, 0.3, 0.1, 0.2, 0.4, so the largest is 0.4. - Normal: bit/s kbit/s. High: bit/s kbit/s. Normal quality needs half the bit rate, so it can stream without buffering on a connection that cannot sustain the high-quality rate.
- The file size claim is correct (size is proportional to rate × resolution, and ). The quality claim is not: the two settings affect different aspects. More samples capture higher frequencies and the shape of the wave more often, but fewer bits per sample give fewer amplitude levels and larger quantisation errors (more noise). The result will sound different, and for most sounds halving a 16-bit resolution to 8 bits would noticeably reduce quality.