Floating and Fixed Point

Floating and Fixed Point are some of those things that I did not get a good explanation when I was being taught about computers. Which is a problem, because they are very integral to how modern computers operate.

Except it hasn't been like that forever. Because when I was a little egg, computers treated floating point processors (FPUs) kinda like they treated GPUs in the late 1990s/early 2000s. It was something extra that you could add to the system, not a core part of it.

Which makes math in old computers something wild.

Let's Begin With Natural Numbers

In general computers are good at representing integers. That's a lie.

Computers can have registers that can be used to represent a range of integer numbers. Because of historical reasons they tend to come in sizes that are multiples of 256, although this is kind of not they way things had to be. Many old computers operated on 127 number ranges.

How do these numbers fit into registers? In binary.

binary decimal
00000000 0
00000001 1
00000010 2
00000011 3
00000100 4
... ...
11111111 255

Two possible values for each one of eight bits gives use 28=256 values in an 8-bit register. Of course modern computers tend to use much bigger registers, generally 64-bit (yes, that's what the bittage of computers means).

What About Negative Numbers?

If you want to represent the sign in a number you need to put it somewhere and a register only has those eight (or more) bits to work with, meaning that some of those positive numbers will become non-representable.

There have been a bunch of ways in which negative numbers have been represented in registers but the standard has become two's complement:

binary decimal
10000000 −128
... ...
11111100 −4
11111101 −3
11111110 −2
11111111 −1
00000000 0
00000001 1
00000010 2
00000011 3
00000100 4
... ...
01111111 127

Might look a bit weird but it has a huge advantage: with this system, the same circuit that adds or subtracts a number from a positive one works to add or subtract a number from a negative one.

The thing is that we're still representing only 256 possible numbers but this time we have chosen to use the range from −128 to 127. That's an unavoidable fact: we can only represent as many numbers as possible 0 and 1 combinations in the register.

Fixed Point, or Naive Decimal Parts

This is the simplest way of representing a range of… well, we're not even calling them real numbers at this point, but numbers with a decimal part. You just take the integer table from the last section and divide the numbers of the right by a power of two. That means that the last N bits in the byte are representing the decimal part:

binary decimal 1 bit 2 bit 4 bit
10000000 −128 −64 −32 −8
... ... ... ... ...
11111100 −4 −2 −1 −0.25
11111101 −3 −1.5 −0.75 −0.1875
11111110 −2 −1 −0.5 −0.125
11111111 −1 −0.5 −0.25 −0.0625
00000000 0 0 0 0
00000001 1 0.5 0.25 0.0625
00000010 2 1 0.5 0.125
00000011 3 1.5 0.75 0.1875
00000100 4 2 1 0.25
... ... ... ... ...
01111111 127 −63.5 −31.25 −7.9375

So, by assigning some of the bits to the decimal part we can represent a finer collection of numbers, but at the cost of having less total range to work with. It also increases a bit the cost of some operations, because adding or subtracting together fixed-point numbers uses the same circuit as adding or subtracting together integers but multiplication and division requires you to do some bit shifting of the result if you want it to have the decimal point in the same place as the inputs.

Floating Point, or Significant Digits

In engineering there's this concept of significant digits, which basically could be summed as "in every operation and scale, there are parts of the input that are relevant and parts that can be considered noise"

In general terms, and for decimal numbers, many people consider six significant digits to be good enough for most purposes, which means that if your number is one-and-a-bit you need to get at least five decimal digits after the point, but if your number is a-hundred-and-forty-two-and-a-you only need three decimal digits after the point.

Which is why there's the exponential notation for numbers, where we just write those significant digits, and then add en "exponent" which just means "move the period by N spaces in one direction":

number exponential
0.00000123456 1.23456e−6
0.0000123456 1.23456e−5
0.000123456 1.23456e−4
0.00123456 1.23456e−3
0.0123456 1.23456e−2
0.123456 1.23456e−1
1.23456 1.23456e0
12.3456 1.23456e1
123.456 1.23456e2
1234.56 1.23456e3
12345.6 1.23456e4
123456. 1.23456e5
1234560. 1.23456e6

The thing is that this format can be used to actually represent binary numbers which are a lot more useful than what fixed point does:

binary decomposed decimal
0001 1000 1×2−8 0.00390625
0001 1111 1×2−1 0.5
0001 0000 1×20 1
0001 0001 1×21 2
0001 0111 1×27 127

This particular distribution of a four-bit base and four-bit exponent gives us as much maximum range as the plain integers, but also allows us to represents numbers as small as the most precise fixed point.

But Wait, There Is a Problem…

Yes. More than one, in fact.

For starters, the total amount of numbers that you can represent are smaller because some numbers can be represented in more than one way:

binary decomposed decimal
0001 0000 1×20 1
0010 1111 2×2−1 1
0100 1110 4×2−2 1
1000 1100 8×2−3 1

This is unavoidable but it is kind of compensated by the actual numbers you can represent being generally more useful, since in a practical context many numbers are not going to be representable.

Then we have the circuitry that you need to make operations with these numbers. It it not trivial. It is in fact much more complex than the one required to do math with integers or fixed-point numbers. Which is the reason that once upon a time FPUs were treated like we treat GPUs now.

And math operations on floating point numbers are prone to rounding errors that can sound very weird to humans. There was an issue once upon a time where a pretty popular family of floating point algorithms resulted in people trying to add 2+2 and getting 3.9 as a result.

But it's been a long time since then, and these days floating point is way beyond the very simple base+exponent format described here. Have a look at the IEEE 754 standard, if you want to know more.

So How Do These Representations Compare?

Here, have a table:

binary unsigned integer fixed (4 bits) floating (4 bits)
0000 0000 0 0 0 0
0000 0001 1 1 0.0625 0
0000 0010 2 2 0.125 0
0000 0011 3 3 0.1875 0
0000 0100 4 4 0.25 0
0000 0101 5 5 0.3125 0
0000 0110 6 6 0.375 0
0000 0111 7 7 0.4375 0
0000 1000 8 8 0.5 0
0000 1001 9 9 0.5625 0
0000 1010 10 10 0.625 0
0000 1011 11 11 0.6875 0
0000 1100 12 12 0.75 0
0000 1101 13 13 0.8125 0
0000 1110 14 14 0.875 0
0000 1111 15 15 0.9375 0
0001 0000 16 16 1 1
0001 0001 17 17 1.0625 2
0001 0010 18 18 1.125 4
0001 0011 19 19 1.1875 8
0001 0100 20 20 1.25 16
0001 0101 21 21 1.3125 32
0001 0110 22 22 1.375 64
0001 0111 23 23 1.4375 128
0001 1000 24 24 1.5 0.00390625
0001 1001 25 25 1.5625 0.0078125
0001 1010 26 26 1.625 0.015625
0001 1011 27 27 1.6875 0.03125
0001 1100 28 28 1.75 0.0625
0001 1101 29 29 1.8125 0.125
0001 1110 30 30 1.875 0.25
0001 1111 31 31 1.9375 0.5
0010 0000 32 32 2 2
0010 0001 33 33 2.0625 4
0010 0010 34 34 2.125 8
0010 0011 35 35 2.1875 16
0010 0100 36 36 2.25 32
0010 0101 37 37 2.3125 64
0010 0110 38 38 2.375 128
0010 0111 39 39 2.4375 256
0010 1000 40 40 2.5 0.0078125
0010 1001 41 41 2.5625 0.015625
0010 1010 42 42 2.625 0.03125
0010 1011 43 43 2.6875 0.0625
0010 1100 44 44 2.75 0.125
0010 1101 45 45 2.8125 0.25
0010 1110 46 46 2.875 0.5
0010 1111 47 47 2.9375 1
0011 0000 48 48 3 3
0011 0001 49 49 3.0625 6
0011 0010 50 50 3.125 12
0011 0011 51 51 3.1875 24
0011 0100 52 52 3.25 48
0011 0101 53 53 3.3125 96
0011 0110 54 54 3.375 192
0011 0111 55 55 3.4375 384
0011 1000 56 56 3.5 0.0117188
0011 1001 57 57 3.5625 0.0234375
0011 1010 58 58 3.625 0.046875
0011 1011 59 59 3.6875 0.09375
0011 1100 60 60 3.75 0.1875
0011 1101 61 61 3.8125 0.375
0011 1110 62 62 3.875 0.75
0011 1111 63 63 3.9375 1.5
0100 0000 64 64 4 4
0100 0001 65 65 4.0625 8
0100 0010 66 66 4.125 16
0100 0011 67 67 4.1875 32
0100 0100 68 68 4.25 64
0100 0101 69 69 4.3125 128
0100 0110 70 70 4.375 256
0100 0111 71 71 4.4375 512
0100 1000 72 72 4.5 0.015625
0100 1001 73 73 4.5625 0.03125
0100 1010 74 74 4.625 0.0625
0100 1011 75 75 4.6875 0.125
0100 1100 76 76 4.75 0.25
0100 1101 77 77 4.8125 0.5
0100 1110 78 78 4.875 1
0100 1111 79 79 4.9375 2
0101 0000 80 80 5 5
0101 0001 81 81 5.0625 10
0101 0010 82 82 5.125 20
0101 0011 83 83 5.1875 40
0101 0100 84 84 5.25 80
0101 0101 85 85 5.3125 160
0101 0110 86 86 5.375 320
0101 0111 87 87 5.4375 640
0101 1000 88 88 5.5 0.0195312
0101 1001 89 89 5.5625 0.0390625
0101 1010 90 90 5.625 0.078125
0101 1011 91 91 5.6875 0.15625
0101 1100 92 92 5.75 0.3125
0101 1101 93 93 5.8125 0.625
0101 1110 94 94 5.875 1.25
0101 1111 95 95 5.9375 2.5
0110 0000 96 96 6 6
0110 0001 97 97 6.0625 12
0110 0010 98 98 6.125 24
0110 0011 99 99 6.1875 48
0110 0100 100 100 6.25 96
0110 0101 101 101 6.3125 192
0110 0110 102 102 6.375 384
0110 0111 103 103 6.4375 768
0110 1000 104 104 6.5 0.0234375
0110 1001 105 105 6.5625 0.046875
0110 1010 106 106 6.625 0.09375
0110 1011 107 107 6.6875 0.1875
0110 1100 108 108 6.75 0.375
0110 1101 109 109 6.8125 0.75
0110 1110 110 110 6.875 1.5
0110 1111 111 111 6.9375 3
0111 0000 112 112 7 7
0111 0001 113 113 7.0625 14
0111 0010 114 114 7.125 28
0111 0011 115 115 7.1875 56
0111 0100 116 116 7.25 112
0111 0101 117 117 7.3125 224
0111 0110 118 118 7.375 448
0111 0111 119 119 7.4375 896
0111 1000 120 120 7.5 0.0273438
0111 1001 121 121 7.5625 0.0546875
0111 1010 122 122 7.625 0.109375
0111 1011 123 123 7.6875 0.21875
0111 1100 124 124 7.75 0.4375
0111 1101 125 125 7.8125 0.875
0111 1110 126 126 7.875 1.75
0111 1111 127 127 7.9375 3.5
1000 0000 128 −128 −8 −8
1000 0001 129 −127 −7.9375 −16
1000 0010 130 −126 −7.875 −32
1000 0011 131 −125 −7.8125 −64
1000 0100 132 −124 −7.75 −128
1000 0101 133 −123 −7.6875 −256
1000 0110 134 −122 −7.625 −512
1000 0111 135 −121 −7.5625 −1024
1000 1000 136 −120 −7.5 −0.03125
1000 1001 137 −119 −7.4375 −0.0625
1000 1010 138 −118 −7.375 −0.125
1000 1011 139 −117 −7.3125 −0.25
1000 1100 140 −116 −7.25 −0.5
1000 1101 141 −115 −7.1875 −1
1000 1110 142 −114 −7.125 −2
1000 1111 143 −113 −7.0625 −4
1001 0000 144 −112 −7 −7
1001 0001 145 −111 −6.9375 −14
1001 0010 146 −110 −6.875 −28
1001 0011 147 −109 −6.8125 −56
1001 0100 148 −108 −6.75 −112
1001 0101 149 −107 −6.6875 −224
1001 0110 150 −106 −6.625 −448
1001 0111 151 −105 −6.5625 −896
1001 1000 152 −104 −6.5 −0.0273438
1001 1001 153 −103 −6.4375 −0.0546875
1001 1010 154 −102 −6.375 −0.109375
1001 1011 155 −101 −6.3125 −0.21875
1001 1100 156 −100 −6.25 −0.4375
1001 1101 157 −99 −6.1875 −0.875
1001 1110 158 −98 −6.125 −1.75
1001 1111 159 −97 −6.0625 −3.5
1010 0000 160 −96 −6 −6
1010 0001 161 −95 −5.9375 −12
1010 0010 162 −94 −5.875 −24
1010 0011 163 −93 −5.8125 −48
1010 0100 164 −92 −5.75 −96
1010 0101 165 −91 −5.6875 −192
1010 0110 166 −90 −5.625 −384
1010 0111 167 −89 −5.5625 −768
1010 1000 168 −88 −5.5 −0.0234375
1010 1001 169 −87 −5.4375 −0.046875
1010 1010 170 −86 −5.375 −0.09375
1010 1011 171 −85 −5.3125 −0.1875
1010 1100 172 −84 −5.25 −0.375
1010 1101 173 −83 −5.1875 −0.75
1010 1110 174 −82 −5.125 −1.5
1010 1111 175 −81 −5.0625 −3
1011 0000 176 −80 −5 −5
1011 0001 177 −79 −4.9375 −10
1011 0010 178 −78 −4.875 −20
1011 0011 179 −77 −4.8125 −40
1011 0100 180 −76 −4.75 −80
1011 0101 181 −75 −4.6875 −160
1011 0110 182 −74 −4.625 −320
1011 0111 183 −73 −4.5625 −640
1011 1000 184 −72 −4.5 −0.0195312
1011 1001 185 −71 −4.4375 −0.0390625
1011 1010 186 −70 −4.375 −0.078125
1011 1011 187 −69 −4.3125 −0.15625
1011 1100 188 −68 −4.25 −0.3125
1011 1101 189 −67 −4.1875 −0.625
1011 1110 190 −66 −4.125 −1.25
1011 1111 191 −65 −4.0625 −2.5
1100 0000 192 −64 −4 −4
1100 0001 193 −63 −3.9375 −8
1100 0010 194 −62 −3.875 −16
1100 0011 195 −61 −3.8125 −32
1100 0100 196 −60 −3.75 −64
1100 0101 197 −59 −3.6875 −128
1100 0110 198 −58 −3.625 −256
1100 0111 199 −57 −3.5625 −512
1100 1000 200 −56 −3.5 −0.015625
1100 1001 201 −55 −3.4375 −0.03125
1100 1010 202 −54 −3.375 −0.0625
1100 1011 203 −53 −3.3125 −0.125
1100 1100 204 −52 −3.25 −0.25
1100 1101 205 −51 −3.1875 −0.5
1100 1110 206 −50 −3.125 −1
1100 1111 207 −49 −3.0625 −2
1101 0000 208 −48 −3 −3
1101 0001 209 −47 −2.9375 −6
1101 0010 210 −46 −2.875 −12
1101 0011 211 −45 −2.8125 −24
1101 0100 212 −44 −2.75 −48
1101 0101 213 −43 −2.6875 −96
1101 0110 214 −42 −2.625 −192
1101 0111 215 −41 −2.5625 −384
1101 1000 216 −40 −2.5 −0.0117188
1101 1001 217 −39 −2.4375 −0.0234375
1101 1010 218 −38 −2.375 −0.046875
1101 1011 219 −37 −2.3125 −0.09375
1101 1100 220 −36 −2.25 −0.1875
1101 1101 221 −35 −2.1875 −0.375
1101 1110 222 −34 −2.125 −0.75
1101 1111 223 −33 −2.0625 −1.5
1110 0000 224 −32 −2 −2
1110 0001 225 −31 −1.9375 −4
1110 0010 226 −30 −1.875 −8
1110 0011 227 −29 −1.8125 −16
1110 0100 228 −28 −1.75 −32
1110 0101 229 −27 −1.6875 −64
1110 0110 230 −26 −1.625 −128
1110 0111 231 −25 −1.5625 −256
1110 1000 232 −24 −1.5 −0.0078125
1110 1001 233 −23 −1.4375 −0.015625
1110 1010 234 −22 −1.375 −0.03125
1110 1011 235 −21 −1.3125 −0.0625
1110 1100 236 −20 −1.25 −0.125
1110 1101 237 −19 −1.1875 −0.25
1110 1110 238 −18 −1.125 −0.5
1110 1111 239 −17 −1.0625 −1
1111 0000 240 −16 −1 −1
1111 0001 241 −15 −0.9375 −2
1111 0010 242 −14 −0.875 −4
1111 0011 243 −13 −0.8125 −8
1111 0100 244 −12 −0.75 −16
1111 0101 245 −11 −0.6875 −32
1111 0110 246 −10 −0.625 −64
1111 0111 247 −9 −0.5625 −128
1111 1000 248 −8 −0.5 −0.00390625
1111 1001 249 −7 −0.4375 −0.0078125
1111 1010 250 −6 −0.375 −0.015625
1111 1011 251 −5 −0.3125 −0.03125
1111 1100 252 −4 −0.25 −0.0625
1111 1101 253 −3 −0.1875 −0.125
1111 1110 254 −2 −0.125 −0.25
1111 1111 255 −1 −0.0625 −0.5