Character sets

AS · 10 min

When you type the letter A, the computer does not store a letter: it stores a number, and a character set says which number means A. This note covers ASCII, extended ASCII and Unicode, how text is held in binary, and why the order of character codes matters for comparing and sorting strings. Paper 1 asks you to describe and compare character sets; Paper 2 relies on the same ideas whenever you use ASC, CHR or compare characters in pseudocode.

What a character set is

Definition

A character set is the complete list of characters that a computer can represent, together with the unique binary code (the character code) assigned to each one.

Every character, including letters, digits, punctuation, the space and invisible control characters such as carriage return or tab, gets a code. Text is stored as the sequence of codes for its characters. For two computers to exchange text correctly, they must use the same character set, which is why standards exist.

The number of bits used per character fixes how many characters the set can hold: nn bits allow 2n2^n different codes.

ASCII

ASCII (American Standard Code for Information Interchange) uses 7 bits per character, giving 27=1282^7 = 128 codes (0 to 127).

  • Codes 0 to 31 (and 127) are control characters: non-printing instructions such as line feed, carriage return and backspace.
  • Code 32 is the space.
  • The rest are the printable English characters: digits, upper- and lower-case letters and punctuation.

In practice each ASCII character is stored in a full byte, with the extra bit set to 0 (or historically used as a parity bit).

The codes are arranged in a very deliberate order:

  • the digits '0' to '9' are consecutive (48 to 57);
  • the upper-case letters 'A' to 'Z' are consecutive (65 to 90);
  • the lower-case letters 'a' to 'z' are consecutive (97 to 122);
  • each lower-case letter is exactly 32 more than its upper-case partner, which in binary means only bit 5 differs (A is 01000001, a is 01100001).

You do not need to memorise any codes: the syllabus says they will be given. But you should know that the groups are consecutive, because questions rely on it. Given that 'A' is 65, you can work out that 'H' is 65+7=7265 + 7 = 72.

Watch out

The character '7' is not the number 7. '7' has ASCII code 55 (00110111); the integer 7 is 00000111. To turn a digit character into its value, subtract the code of '0': 55−48=755 - 48 = 7. Adding the strings "7" and "5" with & gives "75", not 12.

Extended ASCII

Extended ASCII uses all 8 bits of a byte, giving 28=2562^8 = 256 codes. Codes 0 to 127 are identical to ASCII, so plain ASCII text is still valid. Codes 128 to 255 add accented letters (é, ü), currency symbols such as £, and line-drawing symbols.

The problem: there is no single extended ASCII. Different code pages assign different characters to codes 128 to 255 (Western European, Cyrillic, Greek...). A document written with one code page and opened with another shows the wrong symbols, and 256 codes can never cover Chinese, Japanese, Arabic or Hindi, each of which needs thousands of characters.

Unicode

Unicode was created to give every character in every writing system a single, unique code, so that text can be exchanged worldwide without ambiguity. It includes all modern scripts, historic scripts, mathematical symbols and emoji.

  • Each character has a code point, written U+ followed by hex digits: A is U+0041, € is U+20AC.
  • There is room for 1 114 112 code points (U+0000 to U+10FFFF); about 150 000 are currently assigned.
  • The first 128 code points are the same as ASCII, so Unicode is backward compatible with ASCII.

Code points are stored using an encoding:

EncodingBytes per characterNotes
UTF-81 to 4ASCII characters use 1 byte, so English text is the same size as ASCII; the web's standard
UTF-162 or 4Common inside Windows, Java and JavaScript
UTF-324Fixed size, simple to index, but uses the most space

So the string Café €5 (7 characters) takes 10 bytes in UTF-8 (é needs 2, € needs 3), 14 bytes in UTF-16 and 28 bytes in UTF-32.

Key result
Character setBits per characterNumber of charactersKey point
ASCII7128English letters, digits, punctuation, control codes
Extended ASCII8256Adds accented letters and symbols; varies by code page
Unicode8 to 32 (UTF-8, UTF-16, UTF-32)Over a million possibleEvery writing system; first 128 codes match ASCII

Benefits and drawbacks

Unicode's benefit is universality: one standard for all languages, so multilingual documents, websites and messages display correctly everywhere. Its drawback is size: depending on the encoding, each character may need two or four bytes, so files and memory use can be larger than with ASCII. ASCII is compact and simple but limited to English characters.

Comparing and sorting text

Because characters are numbers, a computer compares characters by comparing their codes. Strings are compared character by character from the left until two differ.

This explains some results that surprise students:

  • 'Z' (90) comes before 'a' (97), so "Zebra" < "apple" is TRUE in a straightforward comparison.
  • Digits come before letters, so "42" < "Banana".
  • "10" < "9", because '1' (49) is less than '9' (57) at the first character.

Sorting ["apple", "Banana", "cherry", "Zebra", "42"] by character codes gives "42", "Banana", "Zebra", "apple", "cherry". To sort alphabetically regardless of case, programs convert both strings to the same case before comparing.

Working with codes in pseudocode

The pseudocode guide provides LCASE and UCASE for single characters. Exam papers usually also give ASC(ThisChar : CHAR) RETURNS INTEGER (the character code of a character) and CHR(x : INTEGER) RETURNS CHAR (the character with code x) in the insert. Because the codes of a group are consecutive, you can do arithmetic on them.

Storing a word in ASCII

Given that the ASCII code for 'A' is 65 and for 'a' is 97, show how the word Cat is stored in 8-bit ASCII, in binary and in hexadecimal.

Solution

'C' is two letters after 'A': 65+2=6765 + 2 = 67. 'a' is 97. 't' is 19 letters after 'a': 97+19=11697 + 19 = 116.

CharacterDenaryBinaryHex
C670100001143
a970110000161
t1160111010074

The word is stored as three consecutive bytes: 01000011 01100001 01110100.

Choosing a character set

A company is building a messaging app for users in China, India and France. Explain why Unicode should be used instead of extended ASCII, and state one drawback.

Solution

Extended ASCII has only 256 codes, which cannot represent the thousands of Chinese characters, or Devanagari and other scripts used in India, alongside French accented letters. Its codes 128 to 255 also differ between code pages, so text could display as the wrong characters on another device.

Unicode assigns a unique code point to every character in all these writing systems, so a message displays the same on every device.

Drawback: characters may need more than one byte (for example up to 4 bytes in UTF-8), so messages use more storage and bandwidth than single-byte ASCII.

Converting to upper case using character codes

Write a pseudocode function ToUpper that takes a string and returns it with every lower-case letter changed to upper case. Other characters are unchanged. Do not use UCASE; use ASC and CHR.

Solution

Lower-case letters are 32 more than their upper-case partners, so subtract 32 from the code of any character between 'a' and 'z'.

FUNCTION ToUpper(InString : STRING) RETURNS STRING
   DECLARE OutString : STRING
   DECLARE Index : INTEGER
   DECLARE ThisChar : CHAR

   OutString ← ""
   FOR Index ← 1 TO LENGTH(InString)
      ThisChar ← MID(InString, Index, 1)
      IF ThisChar >= 'a' AND ThisChar <= 'z' THEN
         ThisChar ← CHR(ASC(ThisChar) - 32)
      ENDIF
      OutString ← OutString & ThisChar
   NEXT Index
   RETURN OutString
ENDFUNCTION

ToUpper("Room 101b!") returns "ROOM 101B!": the digits, space and ! are outside the range 'a' to 'z' and pass through unchanged.

A Caesar cipher

A Caesar cipher replaces each upper-case letter with the letter Key places later in the alphabet, wrapping from Z back to A. Other characters are unchanged. Write a pseudocode function Encrypt(PlainText : STRING, Key : INTEGER) RETURNS STRING. State the result of Encrypt("ZOO KEEPER", 5).

Solution

Convert each letter to a position 0 to 25 by subtracting the code of 'A', add the key, use MOD 26 to wrap around, then convert back.

FUNCTION Encrypt(PlainText : STRING, Key : INTEGER) RETURNS STRING
   DECLARE CipherText : STRING
   DECLARE Index : INTEGER
   DECLARE Position : INTEGER
   DECLARE ThisChar : CHAR

   CipherText ← ""
   FOR Index ← 1 TO LENGTH(PlainText)
      ThisChar ← MID(PlainText, Index, 1)
      IF ThisChar >= 'A' AND ThisChar <= 'Z' THEN
         Position ← ASC(ThisChar) - ASC('A')
         Position ← (Position + Key) MOD 26
         ThisChar ← CHR(Position + ASC('A'))
      ENDIF
      CipherText ← CipherText & ThisChar
   NEXT Index
   RETURN CipherText
ENDFUNCTION

For "ZOO KEEPER" with key 5: Z (position 25) becomes (25+5) mod 26=4(25 + 5) \bmod 26 = 4, which is E; O (14) becomes 19, T; K becomes P; E becomes J; P becomes U; R becomes W. The space is unchanged.

Result: "ETT PJJUJW".

Watch out

"Unicode uses 16 bits." This was true of early Unicode, but modern Unicode has over a million code points and is stored in UTF-8, UTF-16 or UTF-32. Say "Unicode can use more bits per character than ASCII, for example up to 32".

Confusing the number of bits with the number of characters. ASCII is 7 bits and 128 characters; extended ASCII is 8 bits and 256 characters. Do not say ASCII has 256 characters.

Forgetting the wrap-around in cipher questions. Without MOD 26, Z plus 5 becomes a punctuation character rather than E.

Exam tip

"Describe" questions on character sets usually want: the number of bits, the number of characters, and what the set includes or cannot include. For "explain why Unicode is needed" give the reason (ASCII cannot represent other languages' characters) and the consequence (Unicode gives a unique code for characters in all languages, so text is shared correctly).

When a question gives you one code (for example 'A' is 65) and asks for another, show the counting step: "'E' is 4 after 'A', so 65+4=6965 + 4 = 69". When writing pseudocode that compares characters, use single quotes for CHAR literals ('a') and double quotes for STRING literals ("a").

Summary
  • A character set maps each character to a unique binary code.
  • ASCII: 7 bits, 128 characters, English only; control codes 0 to 31.
  • Extended ASCII: 8 bits, 256 characters; first 128 match ASCII; upper half varies by code page.
  • Unicode: a unique code point for every character in every writing system; UTF-8, UTF-16, UTF-32 encodings; first 128 code points match ASCII.
  • Letters and digits have consecutive codes; lower case is 32 more than upper case.
  • Strings compare character by character using codes, so "Z" < "a" and "10" < "9".
  • '7' (code 55) is not the integer 7.

Practice

Question
  1. State the number of characters that can be represented in (a) ASCII, (b) extended ASCII.
  2. The ASCII code for 'a' is 97. Give the denary and 8-bit binary codes of 'z' and of 'M'.
  3. Explain why Unicode was developed.
  4. Give one benefit and one drawback of using UTF-32 rather than UTF-8.
  5. A string is stored as the ASCII bytes 01001000 01101001. Given that 'A' is 65 and 'a' is 97, decode the string.
  6. Explain why a program that sorts the names "amir", "Ben" and "Zara" by comparing character codes puts "amir" last.
  7. A student's program reads the digit characters '4' and '9' and joins them with &. Explain why the result is not 13, and describe how the program could obtain the numeric value of a digit character using its code.
  8. Write a pseudocode function CountDigits(Text : STRING) RETURNS INTEGER that returns how many characters in Text are the digits '0' to '9'.
  9. A text message of 160 characters contains only English letters, digits and spaces. Calculate the storage needed in bytes using (a) 8-bit ASCII, (b) UTF-8, (c) UTF-16, (d) UTF-32, and explain why (a) and (b) are equal.
  10. Using the Caesar cipher function from this note, state the output of Encrypt("ATTACK AT DAWN", 13), and explain why applying the same function a second time with key 13 returns the original message.
Answers
  1. (a) 128 (272^7); (b) 256 (282^8).
  2. 'z': 97+25=12297 + 25 = 122 = 01111010. 'M': upper case is 32 less than lower case, so 'A' is 65 and 'M' is 65+12=7765 + 12 = 77 = 01001101.
  3. ASCII and extended ASCII have too few codes (128 or 256) to represent the characters of all the world's languages, and extended ASCII code pages conflict. Unicode gives every character in every writing system a unique code, so text can be exchanged internationally.
  4. Benefit: every character is the same size (4 bytes), so the nnth character can be found directly. Drawback: it uses up to four times as much storage as UTF-8 for English text.
  5. 01001000 = 72 = 'A' + 7 = 'H'. 01101001 = 105 = 'a' + 8 = 'i'. The string is "Hi".
  6. Upper-case codes (65 to 90) are lower than lower-case codes (97 to 122). 'B' (66) and 'Z' (90) are both less than 'a' (97), so "Ben" and "Zara" come before "amir".
  7. & concatenates strings, giving "49". The characters are codes 52 and 57, not the values 4 and 9. Subtract the code of '0': ASC('4') - ASC('0') = 52−48=452 - 48 = 4. (Or use a provided string-to-number function.)
  8. One model answer:
FUNCTION CountDigits(Text : STRING) RETURNS INTEGER
   DECLARE Count : INTEGER
   DECLARE Index : INTEGER
   DECLARE ThisChar : CHAR
   Count ← 0
   FOR Index ← 1 TO LENGTH(Text)
      ThisChar ← MID(Text, Index, 1)
      IF ThisChar >= '0' AND ThisChar <= '9' THEN
         Count ← Count + 1
      ENDIF
   NEXT Index
   RETURN Count
ENDFUNCTION

CountDigits("AB12C9") returns 3.

  1. (a) 160 bytes; (b) 160 bytes; (c) 320 bytes; (d) 640 bytes. UTF-8 stores every ASCII character (code below 128) in a single byte with the same code, so ASCII-only text is identical in UTF-8.
  2. "NGGNPX NG QNJA". Adding 13 twice adds 26, and (p+26) mod 26=p(p + 26) \bmod 26 = p, so every letter returns to its original position.

How well do you know this?

Builds on

Where this leads

Console

Search notes, courses and tools, or run an action