Digital Logic Toolkit — Binary ⇄ text / ASCII
Binary ⇄ Text Converter
Characters to bits and bits back to characters — ASCII and UTF-8, with the code point and byte layout for each one.
Not what you were looking for?
Converting numbers between bases (binary, decimal, hex, octal) → Number base converter. Converting negative numbers → Two’s complement converter.
Text
Binary
Settings
Encoding
Separator
Invalid bytes
Browsers replace invalid bytes with � by default. Reject shows you exactly where the byte stream breaks.
Binary
01001000 01101001
characters (code points): 2 · bytes: 2 · UTF-16 code units: 2
UTF-8 encoding of 2 characters
UTF-8 (RFC 3629): the leading byte says how many bytes follow, and every continuation byte starts 10.
| character | code point | decimal | hex | binary | template |
|---|---|---|---|---|---|
| H | U+0048 | 72 | 48 | 01001000 | 0xxxxxxx |
| i | U+0069 | 105 | 69 | 01101001 | 0xxxxxxx |
| count | value |
|---|---|
| bytes | 2 |
| code points | 2 |
| UTF-16 code units | 2 |
Source: RFC 3629; WHATWG Encoding Standard
Notation used on this page
- Bytes are written most significant bit first.
- The separator between groups of eight is for reading and is not part of the data; a copied bit string has no separators.
- Hex is upper case. Code points are written U+ followed by at least four upper-case hex digits.
Start from a worked example
What binary text actually is
A character is a number. Unicode assigns every character a code point — the letter H is U+0048, decimal 72 — and a character encoding is the rule that turns that code point into bytes. A byte, also called an octet, is eight bits, and a bit is one digit in base 2. So the letter H under ASCII is the single byte 01001000, and under UTF-8 it is the same byte, because UTF-8 was designed to agree with ASCII over the first 128 code points.
There is no such thing as “the” binary for a character until you say which encoding you mean. The character é is one code point, U+00E9, and it is one byte in Latin-1 and two bytes — C3 A9 — in UTF-8. This page always states the encoding it used, and shows the code point, the byte and the eight bits of every byte rather than presenting a lookup you have to trust.
ASCII
ASCII defines 128 code points, U+0000 to U+007F, in four blocks of 32: control codes 0–31, punctuation and digits 32–63, upper-case letters and a few symbols 64–95, and lower-case letters 96–127. Seven bits are enough for all of them, which is why ASCII is a 7-bit code.
The two facts exam questions actually ask for: A = 65 = 01000001 and a = 97 = 01100001. The difference is exactly bit 5 — decimal 32, 0x20 — so case conversion is a single XOR: 65 ⊕ 32 = 97.
The eighth bit of a byte was originally free, and the commonest thing to do with it was to make it a parity bit. That history is on the parity and checksum page: 7-bit ASCII plus one parity bit is a byte.
A printable ASCII table
Every printable ASCII character, from space (32) to tilde (126), with its decimal value, its hexadecimal value and its eight bits. This table is in the page’s HTML rather than drawn by script, so it prints and it is readable without JavaScript.
| Character | Decimal | Hex | Binary |
|---|---|---|---|
| space | 32 | 20 | 00100000 |
| ! | 33 | 21 | 00100001 |
| " | 34 | 22 | 00100010 |
| # | 35 | 23 | 00100011 |
| $ | 36 | 24 | 00100100 |
| % | 37 | 25 | 00100101 |
| & | 38 | 26 | 00100110 |
| ' | 39 | 27 | 00100111 |
| ( | 40 | 28 | 00101000 |
| ) | 41 | 29 | 00101001 |
| * | 42 | 2A | 00101010 |
| + | 43 | 2B | 00101011 |
| , | 44 | 2C | 00101100 |
| - | 45 | 2D | 00101101 |
| . | 46 | 2E | 00101110 |
| / | 47 | 2F | 00101111 |
| 0 | 48 | 30 | 00110000 |
| 1 | 49 | 31 | 00110001 |
| 2 | 50 | 32 | 00110010 |
| 3 | 51 | 33 | 00110011 |
| 4 | 52 | 34 | 00110100 |
| 5 | 53 | 35 | 00110101 |
| 6 | 54 | 36 | 00110110 |
| 7 | 55 | 37 | 00110111 |
| 8 | 56 | 38 | 00111000 |
| 9 | 57 | 39 | 00111001 |
| : | 58 | 3A | 00111010 |
| ; | 59 | 3B | 00111011 |
| < | 60 | 3C | 00111100 |
| = | 61 | 3D | 00111101 |
| > | 62 | 3E | 00111110 |
| ? | 63 | 3F | 00111111 |
| @ | 64 | 40 | 01000000 |
| A | 65 | 41 | 01000001 |
| B | 66 | 42 | 01000010 |
| C | 67 | 43 | 01000011 |
| D | 68 | 44 | 01000100 |
| E | 69 | 45 | 01000101 |
| F | 70 | 46 | 01000110 |
| G | 71 | 47 | 01000111 |
| H | 72 | 48 | 01001000 |
| I | 73 | 49 | 01001001 |
| J | 74 | 4A | 01001010 |
| K | 75 | 4B | 01001011 |
| L | 76 | 4C | 01001100 |
| M | 77 | 4D | 01001101 |
| N | 78 | 4E | 01001110 |
| O | 79 | 4F | 01001111 |
| P | 80 | 50 | 01010000 |
| Q | 81 | 51 | 01010001 |
| R | 82 | 52 | 01010010 |
| S | 83 | 53 | 01010011 |
| T | 84 | 54 | 01010100 |
| U | 85 | 55 | 01010101 |
| V | 86 | 56 | 01010110 |
| W | 87 | 57 | 01010111 |
| X | 88 | 58 | 01011000 |
| Y | 89 | 59 | 01011001 |
| Z | 90 | 5A | 01011010 |
| [ | 91 | 5B | 01011011 |
| \ | 92 | 5C | 01011100 |
| ] | 93 | 5D | 01011101 |
| ^ | 94 | 5E | 01011110 |
| _ | 95 | 5F | 01011111 |
| ` | 96 | 60 | 01100000 |
| a | 97 | 61 | 01100001 |
| b | 98 | 62 | 01100010 |
| c | 99 | 63 | 01100011 |
| d | 100 | 64 | 01100100 |
| e | 101 | 65 | 01100101 |
| f | 102 | 66 | 01100110 |
| g | 103 | 67 | 01100111 |
| h | 104 | 68 | 01101000 |
| i | 105 | 69 | 01101001 |
| j | 106 | 6A | 01101010 |
| k | 107 | 6B | 01101011 |
| l | 108 | 6C | 01101100 |
| m | 109 | 6D | 01101101 |
| n | 110 | 6E | 01101110 |
| o | 111 | 6F | 01101111 |
| p | 112 | 70 | 01110000 |
| q | 113 | 71 | 01110001 |
| r | 114 | 72 | 01110010 |
| s | 115 | 73 | 01110011 |
| t | 116 | 74 | 01110100 |
| u | 117 | 75 | 01110101 |
| v | 118 | 76 | 01110110 |
| w | 119 | 77 | 01110111 |
| x | 120 | 78 | 01111000 |
| y | 121 | 79 | 01111001 |
| z | 122 | 7A | 01111010 |
| { | 123 | 7B | 01111011 |
| | | 124 | 7C | 01111100 |
| } | 125 | 7D | 01111101 |
| ~ | 126 | 7E | 01111110 |
UTF-8
UTF-8 encodes a code point in one to four octets according to its magnitude. The lead octet says how many octets follow, and every continuation octet begins 10.
| Code point range | Octets |
|---|---|
| 0000 0000 – 0000 007F | 0xxxxxxx |
| 0000 0080 – 0000 07FF | 110xxxxx 10xxxxxx |
| 0000 0800 – 0000 FFFF | 1110xxxx 10xxxxxx 10xxxxxx |
| 0001 0000 – 0010 FFFF | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
Two properties follow, and together they are why UTF-8 won. Every ASCII byte is unchanged, so ASCII text is already valid UTF-8. And no continuation octet can be mistaken for a lead octet, so a decoder that starts in the middle of a stream can resynchronise by scanning forward to the next byte that is not 10xxxxxx.
What UTF-8 forbids
- Surrogates. U+D800 to U+DFFF exist only as a UTF-16 mechanism and are not encodable.
- Code points above U+10FFFF. The four-octet form can express more than Unicode defines; the excess is invalid.
- Overlong forms. A code point must use the shortest template that fits it.
RFC 3629 is explicit about why the third rule matters: “a naive implementation may decode the overlong UTF-8 sequence C0 80 into the character U+0000 … Implementations of the decoding algorithm above MUST protect against decoding invalid sequences.” The WHATWG Encoding Standard then fixes what a browser does instead: emit U+FFFD, and “no other behavior is permitted”. This page’s decoder follows that rule by default and offers a strict mode that reports the offset where the byte stream breaks.
Counting: characters, bytes, code units, graphemes
Four different numbers describe the same string, and confusing them is the source of most encoding bugs.
| Text | Code points | UTF-8 bytes | UTF-16 code units | Graphemes |
|---|---|---|---|---|
| Café 😀 | 6 | 10 | 7 | 6 |
| 👨👩👧 | 5 | 18 | 8 | 1 |
"😀".length is 2 in JavaScript, because String.length counts UTF-16 code units, not characters. A field that allows “20 characters” and a database column that allows “20 bytes” are not the same field.
Notation used on this page
- Bytes are written most significant bit first.
- Groups of eight are separated by a space for reading; the space is not part of the data and a copied bit string has none.
- Hexadecimal is upper case.
- Code points are written
U+followed by at least four upper-case hexadecimal digits.
Sources
- RFC 3629 — UTF-8, a transformation format of ISO 10646 (opens in a new tab): the octet templates and the prohibition on overlong forms.
- WHATWG Encoding Standard (opens in a new tab): the decoder state machine and the U+FFFD replacement rule this page implements.
- The Unicode Standard (opens in a new tab), including its best practice for the use of U+FFFD.
Worked examples
- "Hi" → binaryintrotwo bytes, ASCII
- 01001000 01100101 01101100 01101100 01101111 → textintrofive bytes decoded
- "A" → binarycore65, the anchor value
- "a" vs "A" — the case bitcoreone bit apart
- "0" the character vs 0 the numbercorethe most common encoding confusion
- Space and the printable rangecorewhere the control characters end
- "é" in UTF-8coretwo bytes for one character
- An emoji in UTF-8examfour bytes, one character
- "HELLO WORLD" → binaryexam11 bytes including the space
- Binary → text with a corrupted byteedge casewhat to do with input that cannot decode