Floating point, normalisation and character encoding
Floating point needs a stated numerical format; text needs an agreed character encoding. Neither can be decoded safely by guessing what the bits mean.
Content owner: Michael Print · Written for A-Level learners · Checked against official specifications
The idea to start with
Floating point represents a value as mantissa × 2^exponent. Here the mantissa is an eight-bit two's-complement fraction with its binary point immediately after the sign bit; the exponent is a four-bit two's-complement integer. Write these rules before calculating.
Normalisation uses the available significant bits effectively. A nonzero normalised mantissa begins 01 for positive values or 10 for negative values. Character encoding is a separate interpretation: a code point identifies a character, and an encoding specifies its byte representation.
OCR H446 · 1.4.1(g–h,j), with fixed point as a comparison. OCR's clarification uses two's-complement mantissa and exponent; this teaching format is not IEEE 754.
Before you start
Useful foundations
Binary place value and fractions
Two's-complement signed integers
Powers of two with negative exponents
By the end, you should be able to
Decode and normalise positive and negative floating-point values
Align, add or subtract mantissas without changing value
Explain rounding, precision and range
Distinguish character sets, code points and encodings
Read the signed fractional weights
Our eight-bit mantissa has weights −1,1/2,1/4,1/8,1/16,1/32,1/64,1/128. 01101000 therefore means 0.5+0.25+0.0625=0.8125. With exponent 0011 (+3), it represents 0.8125×8=6.5.
Negative 10011000 means −1+0.125+0.0625=−0.8125, giving −6.5 at exponent +3. Inversion and adding one negate the positive mantissa. The leading bit has weight −1; it is not a detached sign followed by magnitude.
Fixed point keeps an agreed binary-point position. Floating point also stores an exponent to span more magnitudes. At a fixed bit budget, mantissa bits improve precision; exponent bits expand exponent range. Four exponent bits here cover −8…+7.
Normalise while preserving value
Normalisation shifts redundant sign bits out while compensating in the exponent. Increasing mantissa magnitude requires decreasing the exponent to preserve value.
Negative 11101000 means −0.1875. Shift left twice to 10100000 (−0.75), reduce exponent to −2 and stop at leading 10.
Zero is a special case: shifting cannot give an all-zero mantissa leading 01 or 10. Our exercise stores zero with an all-zero mantissa and exponent zero.
Mantissa becomes 01100000 = 0.75. Its magnitude has multiplied by four.
3
Compensate in the exponent
Reduce 0 to −2: 1110. The scale factor divides by four.
4
Check
0.75×2^(−2)=0.1875. Leading 01 is normalised in this teaching format.
Align exponents before arithmetic
Use a common exponent, normally the larger one. Shift the smaller-exponent mantissa right by the difference, preserving its sign. Changing only its exponent would change its value.
Add aligned signed mantissas; for subtraction, negate the second mantissa and add. Normalise the result, then round to stored width.
Keep extra working bits where needed. The temporary sum can exceed the mantissa range even when the final normalised value fits. Losing that bit corrupts the answer.
Finite precision makes rounding necessary
Some real values repeat in binary. At exponent zero, this format’s mantissas are spaced by 1/128. Exact 0.7 lies between 89/128=0.6953125 and 90/128=0.703125.
Truncation chooses 0.6953125; nearest rounding chooses 0.703125, 01011010. State the convention when it affects the answer.
Finite precision causes approximation; cancellation can expose fewer significant bits. Exponent overflow means the magnitude exceeds the format’s range. Underflow means a nonzero magnitude is too small. Neither is the same as discarding low mantissa bits during rounding.
ASCII, Unicode and the bytes used for text
ASCII assigns 128 values for English letters, digits, punctuation and controls. Unicode assigns code points across many writing systems. UTF-8, UTF-16 and UTF-32 encode those points with different storage behaviour; Unicode does not require two bytes per character.
In the table, A is U+0041 and a is U+0061; each uses one UTF-8 byte. é is U+00E9 and uses two: C3 A9.
The wrong decoder can turn valid bytes into wrong text. One visible character can also have different Unicode sequences, so character and byte counts differ. Exam questions supply needed values; this table is illustrative.
From a character to bytes
1
Character
The visible character é.
2
Unicode code point
U+00E9 identifies the character in this example.
3
UTF-8 encoding
The byte sequence is C3 A9: two bytes.
Another encoding can store the same code point differently.
Illustrative code points and UTF-8 byte encodings
Character
Code point
UTF-8 hex bytes
Bytes
A
U+0041
41
1
a
U+0061
61
1
é
U+00E9
C3 A9
2
Worked example
Add and subtract 3.25 and 0.5
3.25 is mantissa 01101000 (0.8125), exponent 0010 (+2). 0.5 is mantissa 01000000 (0.5), exponent 0000 (0). Use common exponent +2.
Shift 0.5's mantissa right twice to 00010000 (0.125). Its re-expressed value is 0.125×4=0.5. Add: 01101000 + 00010000 = 01111000, or 0.9375. It is already normalised; 0.9375×4=3.75.
For 3.25−0.5, negate the aligned 0.125 mantissa to 11110000 (−0.125). Add to get 01011000, or 0.6875, at exponent +2. The result is 2.75. Likewise −3.25+0.5 gives 10101000 (−0.6875) at exponent +2, or −2.75.
Subtracting −0.5 from −3.25 adds +0.5, so it uses that same aligned calculation and returns −2.75. Subtracting a negative reverses its sign before addition; it does not subtract the unsigned digits of its stored pattern.
For cancellation, 3.25−3.0 at exponent +2 gives 00001000 (0.0625). Normalise by shifting left three places to 01000000 and decreasing the exponent to −1 (1111): 0.5×0.5=0.25.
All rows use an eight-bit signed fractional mantissa and four-bit signed exponent
Value
Mantissa
Exponent
Check
3.25
01101000
0010
0.8125 × 4
0.5, aligned
00010000
0010
0.125 × 4
3.75
01111000
0010
0.9375 × 4
2.75
01011000
0010
0.6875 × 4
−2.75
10101000
0010
−0.6875 × 4
0.25, normalised
01000000
1111
0.5 × 0.5
Original A-Level practice
6 original questions total 16 marks. Attempt each before opening the independently written indicative marking guidance.
Question 1
3 marks
Decode mantissa 10110000 and exponent 0010 using the format below. [3 marks]
Floating-point format
Value = mantissa × 2^exponent. The mantissa is an eight-bit two's-complement fraction with its binary point immediately after the sign bit. The exponent is a four-bit two's-complement integer. This is the guide's teaching format, rather than IEEE 754.
Show solution and marking guidance+
Indicative answer
The mantissa is −1+1/4+1/8=−0.625 (1). The exponent is +2 (1). Value=−0.625×4=−2.5 (1).
Question 2
3 marks
Normalise mantissa 00011000 with exponent 0000, showing the compensating change. [3 marks]
Floating-point format
Value = mantissa × 2^exponent. The mantissa is an eight-bit two's-complement fraction with its binary point immediately after the sign bit. The exponent is a four-bit two's-complement integer. Preserve the represented value when changing the mantissa and exponent.
Show solution and marking guidance+
Indicative answer
Shift left twice to 01100000 (1). Reduce the exponent from zero to −2, 1110 (1). 0.75×2^(−2)=0.1875 confirms the same value (1).
Question 3
2 marks
Why must the mantissa for 0.5 be shifted right twice before adding it to 3.25 at exponent +2? [2 marks]
Operands before alignment
Value = mantissa × 2^exponent. Each mantissa is an eight-bit two's-complement fraction with its binary point after the sign bit; each exponent is a four-bit two's-complement integer.
Original stored values
Value
Mantissa
Exponent
3.25
01101000
0010
0.5
01000000
0000
Show solution and marking guidance+
Indicative answer
Its exponent must increase from zero to +2 to match the other operand (1). Dividing its mantissa by four compensates, preserving 0.5 as 0.125×4 (1).
Question 4
3 marks
Using the character table below, give the UTF-8 hex bytes and byte length for Aé. [3 marks]
Character encoding input
UTF-8 encodings needed for this question
Character
UTF-8 hex bytes
A
41
é
C3 A9
Show solution and marking guidance+
Indicative answer
A contributes 41 (1); é contributes C3 A9 (1). The resulting sequence 41 C3 A9 has three bytes (1).
Question 5
2 marks
A fixed-width floating-point format gets two extra mantissa bits but keeps the same exponent width. State the main benefit and what does not expand. [2 marks]
Show solution and marking guidance+
Indicative answer
More mantissa bits allow finer precision/more significant binary digits (1). The available exponent range does not expand (1), although exact representable endpoints can change slightly with the mantissa.
Question 6
3 marks
Calculate −3.25−(−0.5) in the format below, using common exponent +2. State the aligned second operand, the mantissa operation and the final value. [3 marks]
Floating-point format
Value = mantissa × 2^exponent. The mantissa is an eight-bit two's-complement fraction with its binary point immediately after the sign bit. The exponent is a four-bit two's-complement integer. Use extra working bits if needed; both operands and the result are exactly representable.
Show solution and marking guidance+
Indicative answer
Aligned −0.5 is mantissa 11110000 (−0.125) with exponent 0010 (1). Negate it to 00010000 and add to −3.25's 10011000, giving 10101000 (1). That represents −0.6875×4=−2.75 (1).
Specification and references
This guide addresses OCR H446 1.4.1(g–h,j), with fixed point as a comparison. OCR's clarification uses two's-complement mantissa and exponent; this teaching format is not IEEE 754.. Check your examination year and the complete specification for the assessment scope.
These are independently written explanations and practice questions. CompSciTutoring.co.uk is not affiliated with or endorsed by an examination board. The marking guidance is indicative; always check the syllabus for your examination year.