Floating-point numbers

A2 · 14 min

Integers in two's complement cannot hold 6.5 or 0.0001, and a fixed binary point cannot hold both very large and very small numbers. Floating-point representation solves this by storing a number as a mantissa (the significant digits) and an exponent (how far to move the binary point), exactly like scientific notation in denary. Floating-point conversion is one of the most reliably examined skills in Paper 3: you will convert in both directions, usually with an 8-, 10- or 12-bit mantissa and a 4- to 8-bit exponent, without a calculator, and explain how the split of bits affects range and precision.

The idea: scientific notation in binary

In denary, 6 370 000=0.637×1076\,370\,000 = 0.637 \times 10^{7}. The digits 0.6370.637 carry the precision; the power 77 carries the size. Binary works the same way with powers of 2:

110.12=0.11012×23110.1_2 = 0.1101_2 \times 2^{3}

Move the binary point three places left and record "3" to say how to put it back.

Definition

A floating-point number is stored in two parts: a mantissa MM, a fixed-point binary fraction holding the significant bits of the number, and an exponent EE, an integer saying how many places to move the binary point. The value is M×2EM \times 2^{E}.

The Cambridge format

In the format Cambridge uses:

  • The mantissa is in two's complement, with the binary point between the first and second bits. The first bit therefore has place value −1-1, and the following bits are 12,14,18,…\tfrac12, \tfrac14, \tfrac18, \dots
  • The exponent is a two's complement integer.
  • The number of bits in each is given in the question, for example "an 8-bit mantissa followed by a 4-bit exponent".
Key result

For an 8-bit mantissa the place values are:

Bit1st2nd3rd4th5th6th7th8th
Place value−1-112\tfrac{1}{2}14\tfrac{1}{4}18\tfrac{1}{8}116\tfrac{1}{16}132\tfrac{1}{32}164\tfrac{1}{64}1128\tfrac{1}{128}

For a 4-bit exponent the place values are −8,4,2,1-8, 4, 2, 1, so the exponent ranges from −8-8 to +7+7.

value=mantissa×2exponent\text{value} = \text{mantissa} \times 2^{\text{exponent}}

So the first bit of the mantissa is the sign: 0 for a positive number, 1 for a negative one. The first bit of the exponent is its sign too: a negative exponent moves the point to the left and makes the number smaller in size.

Converting floating-point binary to denary

There are two reliable ways. Use whichever you make fewer mistakes with, and check with the other.

Floating point to denary: move the point
  1. Convert the exponent from two's complement to denary.
  2. Write the mantissa with its binary point after the first bit.
  3. If the exponent is positive, move the binary point that many places to the right; if negative, that many places to the left, filling new places on the left with copies of the sign bit (0 for positive, 1 for negative).
  4. Convert the resulting two's complement fixed-point number to denary (the leftmost bit carries a negative place value).
Floating point to denary: calculate the mantissa
  1. Convert the exponent from two's complement to denary.
  2. Work out the mantissa as a denary fraction: −1-1 for the first bit if it is 1, plus 12,14,…\tfrac12, \tfrac14, \dots for the other 1 bits.
  3. Multiply the mantissa by 2exponent2^{\text{exponent}}.
A positive number with a positive exponent

A floating-point number has an 8-bit mantissa and a 4-bit exponent, both in two's complement:

Mantissa 01011000, exponent 0011. Find its denary value.

Solution

Exponent: 0011 =3= 3.

Mantissa: 0.1011000, which is 12+18+116=0.6875\tfrac12 + \tfrac18 + \tfrac1{16} = 0.6875.

Value =0.6875×23=5.5= 0.6875 \times 2^{3} = 5.5.

Check by moving the point: 0.1011000 with the point moved 3 places right is 0101.1000, which is 4+1+0.5=5.54 + 1 + 0.5 = 5.5.

A negative mantissa

Mantissa 10100000, exponent 0010 (8-bit mantissa, 4-bit exponent). Find the denary value.

Solution

Exponent: 0010 =2= 2.

Mantissa: 1.0100000 =−1+14=−0.75= -1 + \tfrac14 = -0.75.

Value =−0.75×22=−3= -0.75 \times 2^{2} = -3.

Check by moving the point 2 places right: 101.00000. As a two's complement fixed-point number the place values are −4,2,1-4, 2, 1, so this is −4+1=−3-4 + 1 = -3.

A negative exponent

Mantissa 01100000, exponent 1110. Find the denary value.

Solution

Exponent: 1110 =−8+4+2=−2= -8 + 4 + 2 = -2.

Mantissa: 0.1100000 =12+14=0.75= \tfrac12 + \tfrac14 = 0.75.

Value =0.75×2−2=0.75÷4=0.1875= 0.75 \times 2^{-2} = 0.75 \div 4 = 0.1875.

Moving the point 2 places left (filling with the sign bit 0) gives 0.0011, which is 18+116=0.1875\tfrac18 + \tfrac1{16} = 0.1875.

Converting denary to floating-point binary

Denary to floating point
  1. Ignore the sign for now. Convert the size of the number to fixed-point binary: whole part by repeated division by 2 (or place values), fractional part by repeated multiplication by 2 (or place values).
  2. Move the binary point so the number becomes 0.1... (a 0, the point, then a 1). Count the places moved: moving the point left nn places gives exponent +n+n; moving it right nn places gives exponent −n-n.
  3. Write the mantissa 0.1... and pad with 0s on the right to the full mantissa width.
  4. If the number is negative, take the two's complement of the whole mantissa (invert every bit, then add 1 to the last bit).
  5. Write the exponent as a two's complement integer of the stated width.
  6. Check by converting back.

This method produces a normalised number (one that starts 01 or 10), which is what questions almost always want. Why that matters is covered in normalisation.

A positive number

Represent 12.7512.75 as a normalised floating-point number with an 8-bit mantissa and a 4-bit exponent.

Solution

12=1100212 = 1100_2 and 0.75=0.1120.75 = 0.11_2, so 12.75=1100.11212.75 = 1100.11_2.

Move the point 4 places left: 1100.11=0.110011×241100.11 = 0.110011 \times 2^{4}.

Mantissa, padded to 8 bits: 01100110. Exponent 44: 0100.

MantissaExponent
011001100100

Check: 0.1100112=12+14+132+164=0.7968750.110011_2 = \tfrac12 + \tfrac14 + \tfrac1{32} + \tfrac1{64} = 0.796875, and 0.796875×16=12.750.796875 \times 16 = 12.75.

A negative number

Represent −6.5-6.5 as a normalised floating-point number with an 8-bit mantissa and a 4-bit exponent.

Solution

6.5=110.12=0.1101×236.5 = 110.1_2 = 0.1101 \times 2^{3}.

Positive mantissa: 01101000.

Two's complement for −0.8125-0.8125: invert to 10010111, add 1 to get 10011000.

Exponent 3: 0011.

MantissaExponent
100110000011

Check: 1.0011000 =−1+18+116=−0.8125= -1 + \tfrac18 + \tfrac1{16} = -0.8125, and −0.8125×8=−6.5-0.8125 \times 8 = -6.5.

A small number with a negative exponent

Represent 0.156250.15625 with an 8-bit mantissa and a 4-bit exponent.

Solution

0.15625=18+132=0.0010120.15625 = \tfrac18 + \tfrac1{32} = 0.00101_2.

To make it 0.1... the point must move right 2 places: 0.00101=0.101×2−20.00101 = 0.101 \times 2^{-2}.

Mantissa: 01010000. Exponent −2-2 in 4-bit two's complement: 2=00102 = 0010, invert to 1101, add 1 to get 1110.

MantissaExponent
010100001110

Check: 0.625×2−2=0.156250.625 \times 2^{-2} = 0.15625.

Exam-style: a 12-bit mantissa

A computer uses a 12-bit mantissa and a 4-bit exponent, both in two's complement. Show how −19.375-19.375 is stored as a normalised floating-point number.

Solution

19=10011219 = 10011_2, 0.375=14+18=0.01120.375 = \tfrac14 + \tfrac18 = 0.011_2, so 19.375=10011.011219.375 = 10011.011_2.

Move the point 5 places left: 0.10011011×250.10011011 \times 2^{5}.

Positive mantissa (12 bits): 010011011000.

Two's complement: invert to 101100100111, add 1: 101100101000.

Exponent 5: 0101.

MantissaExponent
1011001010000101

Check: 1.01100101000 =−1+14+18+164+1256=−1+0.25+0.125+0.015625+0.00390625=−0.60546875= -1 + \tfrac14 + \tfrac18 + \tfrac1{64} + \tfrac1{256} = -1 + 0.25 + 0.125 + 0.015625 + 0.00390625 = -0.60546875, and −0.60546875×32=−19.375-0.60546875 \times 32 = -19.375.

Tip

A quick way to find a two's complement: copy the bits from the right up to and including the first 1, then invert every bit to the left of it. 01101000 → copy 1000, invert 0110 to 1001 → 10011000.

Range and precision: how the bits are shared

A fixed total number of bits must be split between mantissa and exponent, and the split is a trade-off.

Definition

Precision is how many significant bits a number can be stored with, which decides how close the stored value can be to the true value. It depends on the number of bits in the mantissa.

Range is the spread between the largest and smallest magnitudes that can be stored. It depends on the number of bits in the exponent.

Key result
  • More mantissa bits, fewer exponent bits: greater precision, smaller range.
  • More exponent bits, fewer mantissa bits: greater range, less precision.

The reason: every extra mantissa bit halves the gap between neighbouring representable values; every extra exponent bit roughly squares the largest power of 2 available.

The extremes for a given format

For a normalised number with an 8-bit mantissa and a 4-bit exponent:

QuantityMantissaExponentValue
Largest positive01111111 (127128\tfrac{127}{128})0111 (77)127128×27=127\tfrac{127}{128} \times 2^{7} = 127
Smallest positive (normalised)01000000 (12\tfrac12)1000 (−8-8)12×2−8=2−9=0.001953125\tfrac12 \times 2^{-8} = 2^{-9} = 0.001953125
Most negative10000000 (−1-1)0111 (77)−1×27=−128-1 \times 2^{7} = -128

Compare a 12-bit total split as 6-bit mantissa and 6-bit exponent. The exponent now goes up to 3131, so the largest positive value is 3132×231≈2.08×109\tfrac{31}{32} \times 2^{31} \approx 2.08 \times 10^{9}, far larger than 127; but the mantissa has only 5 bits after the point, so only about 5 significant bits of any number are kept.

Explaining a change of allocation

A system stores real numbers in 16 bits: a 10-bit mantissa and a 6-bit exponent. A designer proposes changing to a 12-bit mantissa and a 4-bit exponent. Describe the effect.

Solution

The mantissa gains 2 bits, so numbers are stored to greater precision: the gap between representable values is a quarter of what it was, so stored values are closer to the true values and rounding errors are smaller.

The exponent loses 2 bits, so its range falls from −32…31-32 \ldots 31 to −8…7-8 \ldots 7. The range of numbers is therefore much smaller: the largest storable magnitude falls from about 2312^{31} to about 272^{7}, and the smallest positive normalised value rises from 2−332^{-33} to 2−92^{-9}. Values that were storable may now cause overflow or underflow.

The total is still 16 bits, so storage requirements do not change.

A converter in Python

Python's float uses the IEEE 754 standard, which is not the Cambridge format, so you cannot inspect its bits to check answers. These functions implement the Cambridge format exactly and are a good way to check your own conversions when revising.

from fractions import Fraction

def twos_complement_value(bits):
    """Integer value of a two's complement bit string."""
    value = int(bits, 2)
    if bits[0] == "1":
        value -= 1 << len(bits)
    return value

def float_to_denary(mantissa, exponent):
    """Mantissa has its binary point after the first bit."""
    m = Fraction(twos_complement_value(mantissa), 1 << (len(mantissa) - 1))
    e = twos_complement_value(exponent)
    return m * Fraction(2) ** e

def to_bits(value, width):
    """Two's complement bit string of an integer."""
    return format(value & ((1 << width) - 1), f"0{width}b")

def denary_to_float(x, m_bits, e_bits):
    """Normalised representation, truncating any bits that do not fit."""
    x = Fraction(x)
    for e in range(-(1 << (e_bits - 1)), 1 << (e_bits - 1)):
        m = x / Fraction(2) ** e
        if Fraction(1, 2) <= m < 1 or -1 <= m < Fraction(-1, 2):
            scaled = m * (1 << (m_bits - 1))
            whole = scaled.numerator // scaled.denominator
            return to_bits(whole, m_bits), to_bits(e, e_bits)
    raise OverflowError("cannot be represented")

print(float(float_to_denary("01011000", "0011")))   # 5.5
print(float(float_to_denary("10100000", "0010")))   # -3.0
print(float(float_to_denary("01100000", "1110")))   # 0.1875
print(denary_to_float("12.75", 8, 4))               # ('01100110', '0100')
print(denary_to_float("-6.5", 8, 4))                # ('10011000', '0011')
print(denary_to_float("0.15625", 8, 4))             # ('01010000', '1110')
print(denary_to_float("-19.375", 12, 4))            # ('101100101000', '0101')
Watch out
  • The binary point is after the first bit of the mantissa, not before it. 01011000 is 0.101100020.1011000_2, not .010110002.01011000_2.
  • For a negative number, take the two's complement of the mantissa only, after normalising the positive version. Never negate the exponent because the number is negative: the exponent's sign is about the size of the number, not its sign.
  • When moving the point to the left with a negative mantissa, the new leading places are filled with 1s, not 0s.
  • A negative exponent written as sign-and-magnitude (1010 for −2-2) is wrong; the exponent is two's complement (1110).
Exam tip
  • Show working: the fixed-point binary, the number of places moved, the positive mantissa, the inversion and the addition for a negative. Method marks are available even if one bit is wrong.
  • Write the final answer in the boxes or table exactly as wide as asked, with every bit, including trailing zeros.
  • "Calculate the denary value" questions give a mark for the exponent and a mark for the final value. State the exponent in denary as a separate step.
  • In "effect of changing the number of bits" questions, use the words precision (mantissa) and range (exponent), say which increases and which decreases, and mention that the total number of bits is unchanged.
  • No calculators are allowed: practise powers of 2 from 2−102^{-10} to 2102^{10} until they are automatic.
Summary
  • Floating point stores a number as mantissa × 2exponent\times\ 2^{\text{exponent}}.
  • In the Cambridge format, both parts are two's complement; the mantissa's binary point is after its first bit, so its place values are −1,12,14,…-1, \tfrac12, \tfrac14, \dots
  • To decode: convert the exponent, then either move the point or calculate the mantissa and multiply by 2E2^{E}.
  • To encode: convert to fixed-point binary, move the point to get 0.1..., count the moves (left = positive exponent), pad, take the two's complement of the mantissa if negative, write the exponent in two's complement.
  • More mantissa bits give more precision; more exponent bits give more range.
  • For 8-bit mantissa and 4-bit exponent: largest 127127, most negative −128-128, smallest positive normalised 2−92^{-9}.

Practice questions

Question
  1. A floating-point number uses a 10-bit mantissa and a 6-bit exponent, both in two's complement. Convert to denary: (a) mantissa 0110100000, exponent 000011; (b) mantissa 1011000000, exponent 000100; (c) mantissa 0101000000, exponent 111101.
  2. Using an 8-bit mantissa and a 4-bit exponent, represent 27.2527.25 as a normalised floating-point number.
  3. Using an 8-bit mantissa and a 4-bit exponent, represent −10.25-10.25 as a normalised floating-point number.
  4. Using an 8-bit mantissa and a 4-bit exponent, represent 0.02343750.0234375 as a normalised floating-point number.
  5. Using a 12-bit mantissa and a 4-bit exponent, represent −0.0390625-0.0390625 as a normalised floating-point number.
  6. A number is stored with an 8-bit mantissa and a 4-bit exponent. Find (a) the largest positive number, (b) the most negative number, (c) the smallest positive normalised number that can be stored. Give the binary and denary values.
  7. A 16-bit floating-point format can be split as a 12-bit mantissa and 4-bit exponent, or an 8-bit mantissa and 8-bit exponent. A scientist must store distances between stars, which are very large, and needs only about two significant denary figures. Recommend a format and justify it.
  8. Mantissa 101101110000 and exponent 0011 (12-bit and 4-bit two's complement). Find the denary value, showing your method.
Answers
  1. (a) Exponent 33. Mantissa 0.110100000 =12+14+116=0.8125= \tfrac12 + \tfrac14 + \tfrac1{16} = 0.8125. Value 0.8125×8=6.50.8125 \times 8 = 6.5.

    (b) Exponent 44. Mantissa 1.011000000 =−1+14+18=−0.625= -1 + \tfrac14 + \tfrac18 = -0.625. Value −0.625×16=−10-0.625 \times 16 = -10.

    (c) Exponent 111101 =−32+16+8+4+1=−3= -32 + 16 + 8 + 4 + 1 = -3. Mantissa 0.101000000 =0.625= 0.625. Value 0.625×2−3=0.0781250.625 \times 2^{-3} = 0.078125.

  2. 27.25=11011.012=0.1101101×2527.25 = 11011.01_2 = 0.1101101 \times 2^{5}. Mantissa 01101101, exponent 0101. Check: 0.11011012=1091280.1101101_2 = \tfrac{109}{128}; 109128×32=27.25\tfrac{109}{128} \times 32 = 27.25.

  3. 10.25=1010.012=0.101001×2410.25 = 1010.01_2 = 0.101001 \times 2^{4}. Positive mantissa 01010010; two's complement 10101110. Exponent 0100. Check: −1+14+116+132+164=−0.640625-1 + \tfrac14 + \tfrac1{16} + \tfrac1{32} + \tfrac1{64} = -0.640625; ×16=−10.25\times 16 = -10.25.

  4. 0.0234375=3128=0.00000112=0.11×2−50.0234375 = \tfrac{3}{128} = 0.0000011_2 = 0.11 \times 2^{-5}. Mantissa 01100000. Exponent −5-5: 5=01015 = 0101, two's complement 1011. Check: 0.75×2−5=0.02343750.75 \times 2^{-5} = 0.0234375.

  5. 0.0390625=5128=0.00001012=0.101×2−40.0390625 = \tfrac{5}{128} = 0.0000101_2 = 0.101 \times 2^{-4}. Positive mantissa 010100000000; two's complement 101100000000. Exponent −4-4: 1100. Check: (−1+14+18)×2−4=−0.625÷16=−0.0390625(-1 + \tfrac14 + \tfrac18) \times 2^{-4} = -0.625 \div 16 = -0.0390625.

  6. (a) Mantissa 01111111, exponent 0111: 127128×27=127\tfrac{127}{128} \times 2^{7} = 127. (b) Mantissa 10000000, exponent 0111: −1×27=−128-1 \times 2^{7} = -128. (c) Mantissa 01000000, exponent 1000: 12×2−8=2−9=0.001953125\tfrac12 \times 2^{-8} = 2^{-9} = 0.001953125.

  7. The 8-bit mantissa and 8-bit exponent. The exponent then ranges from −128-128 to 127127, so magnitudes up to about 21272^{127} can be stored, which is needed for very large astronomical distances; a 4-bit exponent only reaches 272^{7}, so these values would overflow. An 8-bit mantissa (7 bits after the point) gives precision of roughly 1 part in 128, about two significant denary figures, which is all the scientist needs. The trade-off of less precision for more range is therefore the right one.

  8. Exponent 0011 =3= 3. Mantissa 1.01101110000 =−1+14+18+132+164+1128=−1+0.4296875=−0.5703125= -1 + \tfrac14 + \tfrac18 + \tfrac1{32} + \tfrac1{64} + \tfrac1{128} = -1 + 0.4296875 = -0.5703125. Value =−0.5703125×8=−4.5625= -0.5703125 \times 8 = -4.5625. (Moving the point 3 places right gives 1011.01110000; place values −8+2+1+14+18+116=−4.5625-8 + 2 + 1 + \tfrac14 + \tfrac18 + \tfrac1{16} = -4.5625.)

How well do you know this?

Where this leads

Console

Search notes, courses and tools, or run an action