TTKTheTextKit
Format Conversion

HTML Entity Encoder / Decoder

Switch between Encode and Decode to escape markup characters as HTML entities, or turn entities such as é and € back into the characters they stand for. You can escape only the five markup characters or every non-ASCII character, and the work happens in your browser.

Encoded
<p class="note">Café & crème costs €5 — "fresh" today</p>

Related Tools

[4 suggestions]

About the HTML Entity Encoder / Decoder

An HTML character reference writes a character as text that cannot be mistaken for markup. It starts with an ampersand and ends with a semicolon. A named reference, such as ©, uses a name, and a numeric one uses the code point in decimal, as in ©, or in hex, as in ©. Escaping the five markup characters, the ampersand, less-than, greater-than, double quote and single quote, is what keeps user text from being read as tags or breaking an attribute. This tool can stop there, or also write every non-ASCII character as a reference, using a name where one exists and a number otherwise, which is useful for plain-ASCII files and email. Emoji and other characters beyond the basic plane become a single reference to their code point, not two halves. Decoding handles 155 named references, including the whole Latin-1 range, common punctuation and upper-case spellings such as & and ©, plus any decimal or hex number. Numbers from 128 to 159 follow the HTML rule of reading them as Windows-1252 characters, so € is the euro sign. Zero, surrogate halves and values above U+10FFFF become the replacement character. A reference needs its closing semicolon to be decoded here, and names the tool does not know are left as typed and listed. Encoding is not a substitute for a proper template engine or sanitizer when you build pages from untrusted input.

Which characters must be escaped in HTML?▾

The ampersand and less-than sign must always be escaped in text. Inside an attribute value, the quote that surrounds the value must also be escaped. Escaping all five of & < > " and ' is a safe habit.

What is the difference between &#169; and &#xA9;?▾

Nothing. Both are the copyright sign, written with its code point in decimal (169) or hex (A9). The named form &copy; gives the same character.

Why encode non-ASCII characters at all?▾

Plain-ASCII files, some email systems and older tools can mangle accents and symbols. Writing them as entities keeps the characters intact, though a page served as UTF-8 does not need it.

Does decoding handle entities without a semicolon?▾

No. This tool decodes only references that end in a semicolon, which is the standard form. Browsers accept a few old names without one, but relying on that is fragile.