Floating-point numbers
Integers in two's complement cannot hold 6.5 or 0.0001, and a fixed binary point cannot hold both very large and very small numbers. Floating-point representation solves this by storing a number as a mantissa (the significant digits) and an exponent (how far to move the binary point), exactly like scientific notation in denary. Floating-point conversion is one of the most reliably examined skills in Paper 3: you will convert in both directions, usually with an 8-, 10- or 12-bit mantissa and a 4- to 8-bit exponent, without a calculator, and explain how the split of bits affects range and precision.
The idea: scientific notation in binary
In denary, . The digits carry the precision; the power carries the size. Binary works the same way with powers of 2:
Move the binary point three places left and record "3" to say how to put it back.
A floating-point number is stored in two parts: a mantissa , a fixed-point binary fraction holding the significant bits of the number, and an exponent , an integer saying how many places to move the binary point. The value is .
The Cambridge format
In the format Cambridge uses:
- The mantissa is in two's complement, with the binary point between the first and second bits. The first bit therefore has place value , and the following bits are
- The exponent is a two's complement integer.
- The number of bits in each is given in the question, for example "an 8-bit mantissa followed by a 4-bit exponent".
For an 8-bit mantissa the place values are:
| Bit | 1st | 2nd | 3rd | 4th | 5th | 6th | 7th | 8th |
|---|---|---|---|---|---|---|---|---|
| Place value |
For a 4-bit exponent the place values are , so the exponent ranges from to .
So the first bit of the mantissa is the sign: 0 for a positive number, 1 for a negative one. The first bit of the exponent is its sign too: a negative exponent moves the point to the left and makes the number smaller in size.
Converting floating-point binary to denary
There are two reliable ways. Use whichever you make fewer mistakes with, and check with the other.
- Convert the exponent from two's complement to denary.
- Write the mantissa with its binary point after the first bit.
- If the exponent is positive, move the binary point that many places to the right; if negative, that many places to the left, filling new places on the left with copies of the sign bit (0 for positive, 1 for negative).
- Convert the resulting two's complement fixed-point number to denary (the leftmost bit carries a negative place value).
- Convert the exponent from two's complement to denary.
- Work out the mantissa as a denary fraction: for the first bit if it is 1, plus for the other 1 bits.
- Multiply the mantissa by .
A floating-point number has an 8-bit mantissa and a 4-bit exponent, both in two's complement:
Mantissa 01011000, exponent 0011. Find its denary value.
Solution
Exponent: 0011 .
Mantissa: 0.1011000, which is .
Value .
Check by moving the point: 0.1011000 with the point moved 3 places right is 0101.1000, which is .
Mantissa 10100000, exponent 0010 (8-bit mantissa, 4-bit exponent). Find the denary value.
Solution
Exponent: 0010 .
Mantissa: 1.0100000 .
Value .
Check by moving the point 2 places right: 101.00000. As a two's complement fixed-point number the place values are , so this is .
Mantissa 01100000, exponent 1110. Find the denary value.
Solution
Exponent: 1110 .
Mantissa: 0.1100000 .
Value .
Moving the point 2 places left (filling with the sign bit 0) gives 0.0011, which is .
Converting denary to floating-point binary
- Ignore the sign for now. Convert the size of the number to fixed-point binary: whole part by repeated division by 2 (or place values), fractional part by repeated multiplication by 2 (or place values).
- Move the binary point so the number becomes
0.1...(a 0, the point, then a 1). Count the places moved: moving the point left places gives exponent ; moving it right places gives exponent . - Write the mantissa
0.1...and pad with 0s on the right to the full mantissa width. - If the number is negative, take the two's complement of the whole mantissa (invert every bit, then add 1 to the last bit).
- Write the exponent as a two's complement integer of the stated width.
- Check by converting back.
This method produces a normalised number (one that starts 01 or 10), which is what questions almost always want. Why that matters is covered in normalisation.
Represent as a normalised floating-point number with an 8-bit mantissa and a 4-bit exponent.
Solution
and , so .
Move the point 4 places left: .
Mantissa, padded to 8 bits: 01100110. Exponent : 0100.
| Mantissa | Exponent |
|---|---|
01100110 | 0100 |
Check: , and .
Represent as a normalised floating-point number with an 8-bit mantissa and a 4-bit exponent.
Solution
.
Positive mantissa: 01101000.
Two's complement for : invert to 10010111, add 1 to get 10011000.
Exponent 3: 0011.
| Mantissa | Exponent |
|---|---|
10011000 | 0011 |
Check: 1.0011000 , and .
Represent with an 8-bit mantissa and a 4-bit exponent.
Solution
.
To make it 0.1... the point must move right 2 places: .
Mantissa: 01010000. Exponent in 4-bit two's complement: , invert to 1101, add 1 to get 1110.
| Mantissa | Exponent |
|---|---|
01010000 | 1110 |
Check: .
A computer uses a 12-bit mantissa and a 4-bit exponent, both in two's complement. Show how is stored as a normalised floating-point number.
Solution
, , so .
Move the point 5 places left: .
Positive mantissa (12 bits): 010011011000.
Two's complement: invert to 101100100111, add 1: 101100101000.
Exponent 5: 0101.
| Mantissa | Exponent |
|---|---|
101100101000 | 0101 |
Check: 1.01100101000 , and .
A quick way to find a two's complement: copy the bits from the right up to and including the first 1, then invert every bit to the left of it. 01101000 → copy 1000, invert 0110 to 1001 → 10011000.
Range and precision: how the bits are shared
A fixed total number of bits must be split between mantissa and exponent, and the split is a trade-off.
Precision is how many significant bits a number can be stored with, which decides how close the stored value can be to the true value. It depends on the number of bits in the mantissa.
Range is the spread between the largest and smallest magnitudes that can be stored. It depends on the number of bits in the exponent.
- More mantissa bits, fewer exponent bits: greater precision, smaller range.
- More exponent bits, fewer mantissa bits: greater range, less precision.
The reason: every extra mantissa bit halves the gap between neighbouring representable values; every extra exponent bit roughly squares the largest power of 2 available.
The extremes for a given format
For a normalised number with an 8-bit mantissa and a 4-bit exponent:
| Quantity | Mantissa | Exponent | Value |
|---|---|---|---|
| Largest positive | 01111111 () | 0111 () | |
| Smallest positive (normalised) | 01000000 () | 1000 () | |
| Most negative | 10000000 () | 0111 () |
Compare a 12-bit total split as 6-bit mantissa and 6-bit exponent. The exponent now goes up to , so the largest positive value is , far larger than 127; but the mantissa has only 5 bits after the point, so only about 5 significant bits of any number are kept.
A system stores real numbers in 16 bits: a 10-bit mantissa and a 6-bit exponent. A designer proposes changing to a 12-bit mantissa and a 4-bit exponent. Describe the effect.
Solution
The mantissa gains 2 bits, so numbers are stored to greater precision: the gap between representable values is a quarter of what it was, so stored values are closer to the true values and rounding errors are smaller.
The exponent loses 2 bits, so its range falls from to . The range of numbers is therefore much smaller: the largest storable magnitude falls from about to about , and the smallest positive normalised value rises from to . Values that were storable may now cause overflow or underflow.
The total is still 16 bits, so storage requirements do not change.
A converter in Python
Python's float uses the IEEE 754 standard, which is not the Cambridge format, so you cannot inspect its bits to check answers. These functions implement the Cambridge format exactly and are a good way to check your own conversions when revising.
from fractions import Fraction
def twos_complement_value(bits):
"""Integer value of a two's complement bit string."""
value = int(bits, 2)
if bits[0] == "1":
value -= 1 << len(bits)
return value
def float_to_denary(mantissa, exponent):
"""Mantissa has its binary point after the first bit."""
m = Fraction(twos_complement_value(mantissa), 1 << (len(mantissa) - 1))
e = twos_complement_value(exponent)
return m * Fraction(2) ** e
def to_bits(value, width):
"""Two's complement bit string of an integer."""
return format(value & ((1 << width) - 1), f"0{width}b")
def denary_to_float(x, m_bits, e_bits):
"""Normalised representation, truncating any bits that do not fit."""
x = Fraction(x)
for e in range(-(1 << (e_bits - 1)), 1 << (e_bits - 1)):
m = x / Fraction(2) ** e
if Fraction(1, 2) <= m < 1 or -1 <= m < Fraction(-1, 2):
scaled = m * (1 << (m_bits - 1))
whole = scaled.numerator // scaled.denominator
return to_bits(whole, m_bits), to_bits(e, e_bits)
raise OverflowError("cannot be represented")
print(float(float_to_denary("01011000", "0011"))) # 5.5
print(float(float_to_denary("10100000", "0010"))) # -3.0
print(float(float_to_denary("01100000", "1110"))) # 0.1875
print(denary_to_float("12.75", 8, 4)) # ('01100110', '0100')
print(denary_to_float("-6.5", 8, 4)) # ('10011000', '0011')
print(denary_to_float("0.15625", 8, 4)) # ('01010000', '1110')
print(denary_to_float("-19.375", 12, 4)) # ('101100101000', '0101')
- The binary point is after the first bit of the mantissa, not before it.
01011000is , not . - For a negative number, take the two's complement of the mantissa only, after normalising the positive version. Never negate the exponent because the number is negative: the exponent's sign is about the size of the number, not its sign.
- When moving the point to the left with a negative mantissa, the new leading places are filled with 1s, not 0s.
- A negative exponent written as sign-and-magnitude (
1010for ) is wrong; the exponent is two's complement (1110).
- Show working: the fixed-point binary, the number of places moved, the positive mantissa, the inversion and the addition for a negative. Method marks are available even if one bit is wrong.
- Write the final answer in the boxes or table exactly as wide as asked, with every bit, including trailing zeros.
- "Calculate the denary value" questions give a mark for the exponent and a mark for the final value. State the exponent in denary as a separate step.
- In "effect of changing the number of bits" questions, use the words precision (mantissa) and range (exponent), say which increases and which decreases, and mention that the total number of bits is unchanged.
- No calculators are allowed: practise powers of 2 from to until they are automatic.
- Floating point stores a number as mantissa .
- In the Cambridge format, both parts are two's complement; the mantissa's binary point is after its first bit, so its place values are
- To decode: convert the exponent, then either move the point or calculate the mantissa and multiply by .
- To encode: convert to fixed-point binary, move the point to get
0.1..., count the moves (left = positive exponent), pad, take the two's complement of the mantissa if negative, write the exponent in two's complement. - More mantissa bits give more precision; more exponent bits give more range.
- For 8-bit mantissa and 4-bit exponent: largest , most negative , smallest positive normalised .
Practice questions
- A floating-point number uses a 10-bit mantissa and a 6-bit exponent, both in two's complement. Convert to denary: (a) mantissa
0110100000, exponent000011; (b) mantissa1011000000, exponent000100; (c) mantissa0101000000, exponent111101. - Using an 8-bit mantissa and a 4-bit exponent, represent as a normalised floating-point number.
- Using an 8-bit mantissa and a 4-bit exponent, represent as a normalised floating-point number.
- Using an 8-bit mantissa and a 4-bit exponent, represent as a normalised floating-point number.
- Using a 12-bit mantissa and a 4-bit exponent, represent as a normalised floating-point number.
- A number is stored with an 8-bit mantissa and a 4-bit exponent. Find (a) the largest positive number, (b) the most negative number, (c) the smallest positive normalised number that can be stored. Give the binary and denary values.
- A 16-bit floating-point format can be split as a 12-bit mantissa and 4-bit exponent, or an 8-bit mantissa and 8-bit exponent. A scientist must store distances between stars, which are very large, and needs only about two significant denary figures. Recommend a format and justify it.
- Mantissa
101101110000and exponent0011(12-bit and 4-bit two's complement). Find the denary value, showing your method.
Answers
-
(a) Exponent . Mantissa
0.110100000. Value .(b) Exponent . Mantissa
1.011000000. Value .(c) Exponent
111101. Mantissa0.101000000. Value . -
. Mantissa
01101101, exponent0101. Check: ; . -
. Positive mantissa
01010010; two's complement10101110. Exponent0100. Check: ; . -
. Mantissa
01100000. Exponent : , two's complement1011. Check: . -
. Positive mantissa
010100000000; two's complement101100000000. Exponent :1100. Check: . -
(a) Mantissa
01111111, exponent0111: . (b) Mantissa10000000, exponent0111: . (c) Mantissa01000000, exponent1000: . -
The 8-bit mantissa and 8-bit exponent. The exponent then ranges from to , so magnitudes up to about can be stored, which is needed for very large astronomical distances; a 4-bit exponent only reaches , so these values would overflow. An 8-bit mantissa (7 bits after the point) gives precision of roughly 1 part in 128, about two significant denary figures, which is all the scientist needs. The trade-off of less precision for more range is therefore the right one.
-
Exponent
0011. Mantissa1.01101110000. Value . (Moving the point 3 places right gives1011.01110000; place values .)