Skip to main content

Character Sets

  • ASCII - is the most basic character set. It's of length 0 to 127. 128 characters in total.
  • Extended ASCII - It has the same characters as ASCII plus some additional characters. Its length is 0 to 255. 256 characters in total.
  • Unicode - This is the most extensive character set.
Character set vs Encoding

It's extremely important to understand this difference.

  • Character sets - Allocate a code for each character.
  • Encoding - How the is really stored in memory and sent across the network.
Code points

Unicode is the universal standard. It gives a unique number to all possible characters. This number is called a code point.

The code point is sometimes written as hexadecimal. The actual value is still an integer.

UTF-8 Encoding​

This defines how the Unicode code points are efficiently stored and transmitted. That's why the name Unicode Transformation Format.

Each character is Unicode is an integer. Directly storing and transmitting as integer is expensive because integer is 32 bit. Instead, UTF is used to represent these integers in a variable length byte format.

why ASCII has no special encoding?

ASCII has one but since it's just 128 characters, the encoding is to directly use 7 bits to encode each value.

Since this has a direct mapping, we feel there is no encoding.

Unicode Properties​

Unicode standards also defines properties such as uppercase, lowercase, digit, Latin, etc. for every character.

These properties are then exposed by programming languages in different ways. For example, in Java, the Character class has methods such as isUpperCase(), isLowerCase(), isDigit(), etc.