Character sets
When you type the letter A, the computer does not store a letter: it stores a number, and a character set says which number means A. This note covers ASCII, extended ASCII and Unicode, how text is held in binary, and why the order of character codes matters for comparing and sorting strings. Paper 1 asks you to describe and compare character sets; Paper 2 relies on the same ideas whenever you use ASC, CHR or compare characters in pseudocode.
What a character set is
A character set is the complete list of characters that a computer can represent, together with the unique binary code (the character code) assigned to each one.
Every character, including letters, digits, punctuation, the space and invisible control characters such as carriage return or tab, gets a code. Text is stored as the sequence of codes for its characters. For two computers to exchange text correctly, they must use the same character set, which is why standards exist.
The number of bits used per character fixes how many characters the set can hold: bits allow different codes.
ASCII
ASCII (American Standard Code for Information Interchange) uses 7 bits per character, giving codes (0 to 127).
- Codes 0 to 31 (and 127) are control characters: non-printing instructions such as line feed, carriage return and backspace.
- Code 32 is the space.
- The rest are the printable English characters: digits, upper- and lower-case letters and punctuation.
In practice each ASCII character is stored in a full byte, with the extra bit set to 0 (or historically used as a parity bit).
The codes are arranged in a very deliberate order:
- the digits
'0'to'9'are consecutive (48 to 57); - the upper-case letters
'A'to'Z'are consecutive (65 to 90); - the lower-case letters
'a'to'z'are consecutive (97 to 122); - each lower-case letter is exactly 32 more than its upper-case partner, which in binary means only bit 5 differs (
Ais01000001,ais01100001).
You do not need to memorise any codes: the syllabus says they will be given. But you should know that the groups are consecutive, because questions rely on it. Given that 'A' is 65, you can work out that 'H' is .
The character '7' is not the number 7. '7' has ASCII code 55 (00110111); the integer 7 is 00000111. To turn a digit character into its value, subtract the code of '0': . Adding the strings "7" and "5" with & gives "75", not 12.
Extended ASCII
Extended ASCII uses all 8 bits of a byte, giving codes. Codes 0 to 127 are identical to ASCII, so plain ASCII text is still valid. Codes 128 to 255 add accented letters (é, ü), currency symbols such as £, and line-drawing symbols.
The problem: there is no single extended ASCII. Different code pages assign different characters to codes 128 to 255 (Western European, Cyrillic, Greek...). A document written with one code page and opened with another shows the wrong symbols, and 256 codes can never cover Chinese, Japanese, Arabic or Hindi, each of which needs thousands of characters.
Unicode
Unicode was created to give every character in every writing system a single, unique code, so that text can be exchanged worldwide without ambiguity. It includes all modern scripts, historic scripts, mathematical symbols and emoji.
- Each character has a code point, written
U+followed by hex digits: A isU+0041, € isU+20AC. - There is room for 1 114 112 code points (
U+0000toU+10FFFF); about 150 000 are currently assigned. - The first 128 code points are the same as ASCII, so Unicode is backward compatible with ASCII.
Code points are stored using an encoding:
| Encoding | Bytes per character | Notes |
|---|---|---|
| UTF-8 | 1 to 4 | ASCII characters use 1 byte, so English text is the same size as ASCII; the web's standard |
| UTF-16 | 2 or 4 | Common inside Windows, Java and JavaScript |
| UTF-32 | 4 | Fixed size, simple to index, but uses the most space |
So the string Café €5 (7 characters) takes 10 bytes in UTF-8 (é needs 2, € needs 3), 14 bytes in UTF-16 and 28 bytes in UTF-32.
| Character set | Bits per character | Number of characters | Key point |
|---|---|---|---|
| ASCII | 7 | 128 | English letters, digits, punctuation, control codes |
| Extended ASCII | 8 | 256 | Adds accented letters and symbols; varies by code page |
| Unicode | 8 to 32 (UTF-8, UTF-16, UTF-32) | Over a million possible | Every writing system; first 128 codes match ASCII |
Benefits and drawbacks
Unicode's benefit is universality: one standard for all languages, so multilingual documents, websites and messages display correctly everywhere. Its drawback is size: depending on the encoding, each character may need two or four bytes, so files and memory use can be larger than with ASCII. ASCII is compact and simple but limited to English characters.
Comparing and sorting text
Because characters are numbers, a computer compares characters by comparing their codes. Strings are compared character by character from the left until two differ.
This explains some results that surprise students:
'Z'(90) comes before'a'(97), so"Zebra" < "apple"is TRUE in a straightforward comparison.- Digits come before letters, so
"42" < "Banana". "10" < "9", because'1'(49) is less than'9'(57) at the first character.
Sorting ["apple", "Banana", "cherry", "Zebra", "42"] by character codes gives "42", "Banana", "Zebra", "apple", "cherry". To sort alphabetically regardless of case, programs convert both strings to the same case before comparing.
Working with codes in pseudocode
The pseudocode guide provides LCASE and UCASE for single characters. Exam papers usually also give ASC(ThisChar : CHAR) RETURNS INTEGER (the character code of a character) and CHR(x : INTEGER) RETURNS CHAR (the character with code x) in the insert. Because the codes of a group are consecutive, you can do arithmetic on them.
Given that the ASCII code for 'A' is 65 and for 'a' is 97, show how the word Cat is stored in 8-bit ASCII, in binary and in hexadecimal.
Solution
'C' is two letters after 'A': . 'a' is 97. 't' is 19 letters after 'a': .
| Character | Denary | Binary | Hex |
|---|---|---|---|
| C | 67 | 01000011 | 43 |
| a | 97 | 01100001 | 61 |
| t | 116 | 01110100 | 74 |
The word is stored as three consecutive bytes: 01000011 01100001 01110100.
A company is building a messaging app for users in China, India and France. Explain why Unicode should be used instead of extended ASCII, and state one drawback.
Solution
Extended ASCII has only 256 codes, which cannot represent the thousands of Chinese characters, or Devanagari and other scripts used in India, alongside French accented letters. Its codes 128 to 255 also differ between code pages, so text could display as the wrong characters on another device.
Unicode assigns a unique code point to every character in all these writing systems, so a message displays the same on every device.
Drawback: characters may need more than one byte (for example up to 4 bytes in UTF-8), so messages use more storage and bandwidth than single-byte ASCII.
Write a pseudocode function ToUpper that takes a string and returns it with every lower-case letter changed to upper case. Other characters are unchanged. Do not use UCASE; use ASC and CHR.
Solution
Lower-case letters are 32 more than their upper-case partners, so subtract 32 from the code of any character between 'a' and 'z'.
FUNCTION ToUpper(InString : STRING) RETURNS STRING
DECLARE OutString : STRING
DECLARE Index : INTEGER
DECLARE ThisChar : CHAR
OutString ← ""
FOR Index ← 1 TO LENGTH(InString)
ThisChar ← MID(InString, Index, 1)
IF ThisChar >= 'a' AND ThisChar <= 'z' THEN
ThisChar ← CHR(ASC(ThisChar) - 32)
ENDIF
OutString ← OutString & ThisChar
NEXT Index
RETURN OutString
ENDFUNCTIONToUpper("Room 101b!") returns "ROOM 101B!": the digits, space and ! are outside the range 'a' to 'z' and pass through unchanged.
A Caesar cipher replaces each upper-case letter with the letter Key places later in the alphabet, wrapping from Z back to A. Other characters are unchanged. Write a pseudocode function Encrypt(PlainText : STRING, Key : INTEGER) RETURNS STRING. State the result of Encrypt("ZOO KEEPER", 5).
Solution
Convert each letter to a position 0 to 25 by subtracting the code of 'A', add the key, use MOD 26 to wrap around, then convert back.
FUNCTION Encrypt(PlainText : STRING, Key : INTEGER) RETURNS STRING
DECLARE CipherText : STRING
DECLARE Index : INTEGER
DECLARE Position : INTEGER
DECLARE ThisChar : CHAR
CipherText ← ""
FOR Index ← 1 TO LENGTH(PlainText)
ThisChar ← MID(PlainText, Index, 1)
IF ThisChar >= 'A' AND ThisChar <= 'Z' THEN
Position ← ASC(ThisChar) - ASC('A')
Position ← (Position + Key) MOD 26
ThisChar ← CHR(Position + ASC('A'))
ENDIF
CipherText ← CipherText & ThisChar
NEXT Index
RETURN CipherText
ENDFUNCTIONFor "ZOO KEEPER" with key 5: Z (position 25) becomes , which is E; O (14) becomes 19, T; K becomes P; E becomes J; P becomes U; R becomes W. The space is unchanged.
Result: "ETT PJJUJW".
"Unicode uses 16 bits." This was true of early Unicode, but modern Unicode has over a million code points and is stored in UTF-8, UTF-16 or UTF-32. Say "Unicode can use more bits per character than ASCII, for example up to 32".
Confusing the number of bits with the number of characters. ASCII is 7 bits and 128 characters; extended ASCII is 8 bits and 256 characters. Do not say ASCII has 256 characters.
Forgetting the wrap-around in cipher questions. Without MOD 26, Z plus 5 becomes a punctuation character rather than E.
"Describe" questions on character sets usually want: the number of bits, the number of characters, and what the set includes or cannot include. For "explain why Unicode is needed" give the reason (ASCII cannot represent other languages' characters) and the consequence (Unicode gives a unique code for characters in all languages, so text is shared correctly).
When a question gives you one code (for example 'A' is 65) and asks for another, show the counting step: "'E' is 4 after 'A', so ". When writing pseudocode that compares characters, use single quotes for CHAR literals ('a') and double quotes for STRING literals ("a").
- A character set maps each character to a unique binary code.
- ASCII: 7 bits, 128 characters, English only; control codes 0 to 31.
- Extended ASCII: 8 bits, 256 characters; first 128 match ASCII; upper half varies by code page.
- Unicode: a unique code point for every character in every writing system; UTF-8, UTF-16, UTF-32 encodings; first 128 code points match ASCII.
- Letters and digits have consecutive codes; lower case is 32 more than upper case.
- Strings compare character by character using codes, so
"Z" < "a"and"10" < "9". '7'(code 55) is not the integer 7.
Practice
- State the number of characters that can be represented in (a) ASCII, (b) extended ASCII.
- The ASCII code for
'a'is 97. Give the denary and 8-bit binary codes of'z'and of'M'. - Explain why Unicode was developed.
- Give one benefit and one drawback of using UTF-32 rather than UTF-8.
- A string is stored as the ASCII bytes
01001000 01101001. Given that'A'is 65 and'a'is 97, decode the string. - Explain why a program that sorts the names
"amir","Ben"and"Zara"by comparing character codes puts"amir"last. - A student's program reads the digit characters
'4'and'9'and joins them with&. Explain why the result is not 13, and describe how the program could obtain the numeric value of a digit character using its code. - Write a pseudocode function
CountDigits(Text : STRING) RETURNS INTEGERthat returns how many characters inTextare the digits'0'to'9'. - A text message of 160 characters contains only English letters, digits and spaces. Calculate the storage needed in bytes using (a) 8-bit ASCII, (b) UTF-8, (c) UTF-16, (d) UTF-32, and explain why (a) and (b) are equal.
- Using the Caesar cipher function from this note, state the output of
Encrypt("ATTACK AT DAWN", 13), and explain why applying the same function a second time with key 13 returns the original message.
Answers
- (a) 128 (); (b) 256 ().
'z': =01111010.'M': upper case is 32 less than lower case, so'A'is 65 and'M'is =01001101.- ASCII and extended ASCII have too few codes (128 or 256) to represent the characters of all the world's languages, and extended ASCII code pages conflict. Unicode gives every character in every writing system a unique code, so text can be exchanged internationally.
- Benefit: every character is the same size (4 bytes), so the th character can be found directly. Drawback: it uses up to four times as much storage as UTF-8 for English text.
01001000= 72 ='A'+ 7 ='H'.01101001= 105 ='a'+ 8 ='i'. The string is"Hi".- Upper-case codes (65 to 90) are lower than lower-case codes (97 to 122).
'B'(66) and'Z'(90) are both less than'a'(97), so"Ben"and"Zara"come before"amir". &concatenates strings, giving"49". The characters are codes 52 and 57, not the values 4 and 9. Subtract the code of'0':ASC('4') - ASC('0')= . (Or use a provided string-to-number function.)- One model answer:
FUNCTION CountDigits(Text : STRING) RETURNS INTEGER
DECLARE Count : INTEGER
DECLARE Index : INTEGER
DECLARE ThisChar : CHAR
Count ← 0
FOR Index ← 1 TO LENGTH(Text)
ThisChar ← MID(Text, Index, 1)
IF ThisChar >= '0' AND ThisChar <= '9' THEN
Count ← Count + 1
ENDIF
NEXT Index
RETURN Count
ENDFUNCTIONCountDigits("AB12C9") returns 3.
- (a) 160 bytes; (b) 160 bytes; (c) 320 bytes; (d) 640 bytes. UTF-8 stores every ASCII character (code below 128) in a single byte with the same code, so ASCII-only text is identical in UTF-8.
"NGGNPX NG QNJA". Adding 13 twice adds 26, and , so every letter returns to its original position.