Sound representation

AS · 9 min

Sound in the real world is a continuously varying pressure wave; a computer can only store a finite list of binary numbers. Digital sound bridges that gap by sampling: measuring the wave at regular moments and storing each measurement as an integer. This note covers the key terms (sampling, sampling rate, sampling resolution, analogue and digital data), how to calculate the size of a sound file, and how changing the sampling rate and resolution affects accuracy and file size, which Paper 1 tests in nearly every series.

Analogue and digital data

Definition

Analogue data is data that varies continuously: between any two values there are infinitely many others. A sound wave, a temperature or a voltage from a microphone are analogue.

Digital data is data that takes only discrete (separate) values, stored as binary numbers.

A microphone converts the pressure wave into an analogue electrical signal whose voltage rises and falls with the sound. An analogue-to-digital converter (ADC) then samples that voltage and outputs binary numbers. When the sound is played back, a digital-to-analogue converter (DAC) turns the stored numbers back into a varying voltage that drives a loudspeaker.

Sampling

Definition

Sampling is measuring the amplitude (height) of a sound wave at regular time intervals and recording each measurement as a binary number.

The sampling rate is the number of samples taken per second, measured in hertz (Hz).

The sampling resolution is the number of bits used to store each sample (also called the bit depth).

The sampling resolution decides how many different amplitude levels are available. With nn bits there are 2n2^n levels. Each measurement is rounded to the nearest available level, a process called quantisation. The difference between the true amplitude and the stored level is the quantisation error.

The graph shows a sound wave sampled once per time unit with a 3-bit sampling resolution, so there are 23=82^3 = 8 levels, 0 to 7. Each dot is the stored, rounded value.

y = 3.3 + 3 sin(0.6 x) (0, 3) (1, 5) (2, 6) (3, 6) (4, 5) (5, 4) (6, 2) (7, 1) (8, 0) (9, 1) (10, 2)
Time012345678910
True amplitude3.304.996.106.225.333.721.970.690.310.982.46
Stored level35665421012
Binary011101110110101100010001000001010

When this is played back, the output jumps between the stored levels. The digital version is a stepped approximation of the original: it misses what happens between samples, and each sample is slightly wrong because of rounding.

Improving the accuracy of a recording

There are two ways to make the digital version closer to the original wave.

Increase the sampling rate. Taking more samples per second means the wave is measured at more points, so faster changes (higher frequencies) are captured and the playback follows the original shape more closely.

Increase the sampling resolution. Using more bits per sample gives more amplitude levels, so each sample can be stored closer to its true value. Quantisation error is smaller and the range between the quietest and loudest sounds that can be represented (the dynamic range) is greater.

Both improvements cost space: the file size is directly proportional to each.

Extension: the Nyquist rate

Not required by the syllabus, but it explains the numbers you see. To capture a frequency ff, you must sample at more than 2f2f. Human hearing reaches about 20 kHz, so CD audio is sampled at 44.1 kHz. Telephone speech only needs frequencies up to about 4 kHz, so phone systems sample at 8 kHz.

Sound file size

Every second of sound contains (sampling rate) samples, each of (sampling resolution) bits, for each channel. Stereo has two channels (left and right), so it doubles the data.

Key result
file size (bits)=sampling rate (Hz)×sampling resolution (bits)×length (s)×number of channels\text{file size (bits)} = \text{sampling rate (Hz)} \times \text{sampling resolution (bits)} \times \text{length (s)} \times \text{number of channels}

Divide by 8 for bytes. As with images, this is an estimate: a real file also has a header (and may be compressed).

ChangeEffect on file sizeEffect on sound
Double the sampling rateDoublesMore accurate; higher frequencies captured
Increase resolution 8 → 16 bitsDoublesMore levels; less quantisation error; greater dynamic range
Mono → stereoDoublesSeparate left and right channels
Halve the lengthHalvesShorter recording
Calculating a sound file size
  1. Convert the sampling rate to Hz (44.1 kHz = 44 100 Hz) and the length to seconds.
  2. Multiply sampling rate × sampling resolution × time × channels to get bits.
  3. Divide by 8 to get bytes, then convert to the unit asked for.
  4. State the unit and note that the header is ignored.
CD-quality audio

A song lasts 3 minutes. It is recorded in stereo at a sampling rate of 44.1 kHz with a sampling resolution of 16 bits. Estimate the file size in MB.

Solution

Time: 3×60=1803 \times 60 = 180 s. Sampling rate: 44 10044\ 100 Hz. Channels: 2.

Bits: 44 100×16×180×2=254 016 00044\ 100 \times 16 \times 180 \times 2 = 254\ 016\ 000.

Bytes: 254 016 000÷8=31 752 000254\ 016\ 000 \div 8 = 31\ 752\ 000.

Size: 31 752 000÷106≈31.8 MB31\ 752\ 000 \div 10^{6} \approx 31.8\ \text{MB} (about 30.3 MiB30.3\ \text{MiB}).

Speech recording

A voice memo is recorded in mono at 8 kHz with 8-bit samples. Calculate the size of a one-minute memo in kB.

Solution

8000×8×60×1=3 840 0008000 \times 8 \times 60 \times 1 = 3\ 840\ 000 bits =480 000= 480\ 000 bytes =480 kB= 480\ \text{kB}.

(8-bit samples are 1 byte each, so this is simply 8000×608000 \times 60 bytes.)

Working backwards

An audio CD holds 700 MB. CD audio uses 44.1 kHz, 16-bit stereo. Estimate the maximum playing time in minutes.

Solution

Bytes per second: 44 100×16×2÷8=44 100×4=176 40044\ 100 \times 16 \times 2 \div 8 = 44\ 100 \times 4 = 176\ 400 bytes per second.

Seconds: 700×106÷176 400≈3968700 \times 10^6 \div 176\ 400 \approx 3968 s.

Minutes: 3968÷60≈663968 \div 60 \approx 66 minutes.

Explaining the effect of a change

A podcast is recorded in mono at 16 kHz with 16-bit resolution. The producer considers changing to 48 kHz with 24-bit resolution. Calculate the factor by which the file size increases, and explain the effect on the recording.

Solution

File size is proportional to sampling rate × resolution.

Factor =48 000×2416 000×16=3×1.5=4.5= \dfrac{48\ 000 \times 24}{16\ 000 \times 16} = 3 \times 1.5 = 4.5. The file becomes 4.5 times larger.

Effect: three times as many samples per second, so the wave is measured more often and higher frequencies are captured; 24 bits give 2242^{24} levels rather than 2162^{16}, so each sample is closer to the true amplitude (smaller quantisation error). The recording is a more accurate representation of the original sound.

Watch out

Mixing up rate and resolution. Sampling rate is how often (samples per second). Sampling resolution is how precisely (bits per sample). Examiners penalise answers that swap them.

Forgetting stereo. If the question says stereo, multiply by 2.

Forgetting kHz. 44.1 kHz is 44 100 samples per second, not 44.1.

Saying a higher sampling rate makes sound "louder" or "clearer". Use precise language: a higher sampling rate gives a more accurate representation of the wave because it is measured more often; higher resolution reduces quantisation error.

Exam tip

Definitions are worth one mark each and examiners look for key words: sampling rate is "the number of samples taken per second"; sampling resolution is "the number of bits used to store each sample".

For "explain the impact of increasing the sampling resolution", give both the accuracy effect (more amplitude values available, each sample closer to the original, less quantisation error) and the file size effect (more bits per sample, larger file). Many candidates give only one side and lose a mark.

In calculations, write the full expression first, then simplify. No calculator is allowed: cancel factors of 8 early (16 bits = 2 bytes).

Summary
  • Sound is analogue; it is stored digitally by sampling its amplitude at regular intervals.
  • Sampling rate: samples per second (Hz). Sampling resolution: bits per sample.
  • nn-bit resolution gives 2n2^n amplitude levels; rounding to a level is quantisation.
  • Higher rate or resolution gives a more accurate recording and a bigger file.
  • File size (bits) = rate × resolution × seconds × channels; divide by 8 for bytes.
  • An ADC converts analogue to digital; a DAC converts back for playback.

Practice

Question
  1. Explain the difference between analogue data and digital data.
  2. Define the terms sampling rate and sampling resolution.
  3. Calculate the number of amplitude levels available with a sampling resolution of 10 bits.
  4. A 10-second sound effect is recorded in mono at 22 050 Hz with 8-bit samples. Calculate the file size in bytes.
  5. A recording studio records at 48 kHz, 24-bit stereo. Calculate the size of a one-minute recording in MB.
  6. Describe two effects of reducing the sampling rate of a recording.
  7. A 5-minute interview is recorded in mono at 16 kHz with a 16-bit resolution. Calculate the file size in MiB, to one decimal place.
  8. A sound is recorded with a 2-bit sampling resolution. The true amplitudes at five sample points, on a scale where the levels are 0, 1, 2 and 3, are 0.4, 1.7, 2.9, 2.2 and 0.6. State the stored binary values and the largest quantisation error.
  9. A music streaming service offers "normal" quality (sampled at 22.05 kHz, 16-bit stereo) and "high" quality (44.1 kHz, 16-bit stereo), both uncompressed. Calculate the bit rate of each in kilobits per second, and explain why a user with a slow connection might choose normal quality.
  10. A student claims: "If I double the sampling rate and halve the sampling resolution, the file size stays the same, so the sound quality will be the same too." Evaluate this claim.
Answers
  1. Analogue data varies continuously and can take any value in a range; digital data takes only discrete values, stored as binary numbers.
  2. Sampling rate: the number of samples taken per second. Sampling resolution: the number of bits used to store each sample.
  3. 210=10242^{10} = 1024 levels.
  4. 22 050×8×10÷8=220 50022\ 050 \times 8 \times 10 \div 8 = 220\ 500 bytes.
  5. 48 000×24×60×2÷8=17 280 00048\ 000 \times 24 \times 60 \times 2 \div 8 = 17\ 280\ 000 bytes ≈17.3 MB\approx 17.3\ \text{MB}.
  6. Fewer samples per second, so the wave is measured less often: the recording is a less accurate representation and high-frequency sounds may be lost. The file size is smaller (in proportion to the rate), so it needs less storage and bandwidth.
  7. 16 000×16×300÷8=9 600 00016\ 000 \times 16 \times 300 \div 8 = 9\ 600\ 000 bytes; ÷220≈9.2 MiB\div 2^{20} \approx 9.2\ \text{MiB}.
  8. Rounded levels 0, 2, 3, 2, 1, stored as 00, 10, 11, 10, 01. Errors are 0.4, 0.3, 0.1, 0.2, 0.4, so the largest is 0.4.
  9. Normal: 22 050×16×2=705 60022\ 050 \times 16 \times 2 = 705\ 600 bit/s ≈705.6\approx 705.6 kbit/s. High: 44 100×16×2=1 411 20044\ 100 \times 16 \times 2 = 1\ 411\ 200 bit/s ≈1411.2\approx 1411.2 kbit/s. Normal quality needs half the bit rate, so it can stream without buffering on a connection that cannot sustain the high-quality rate.
  10. The file size claim is correct (size is proportional to rate × resolution, and 2×12=12 \times \tfrac12 = 1). The quality claim is not: the two settings affect different aspects. More samples capture higher frequencies and the shape of the wave more often, but fewer bits per sample give fewer amplitude levels and larger quantisation errors (more noise). The result will sound different, and for most sounds halving a 16-bit resolution to 8 bits would noticeably reduce quality.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action