How to Encode URL Special Characters
Replace each special character with a percent sign followed by its two-digit hexadecimal ASCII code — a space becomes %20, an ampersand becomes %26, a question mark becomes %3F. Letters, digits, and - _ . ~ are never encoded. ToolNest's free URL encoder applies percent-encoding to any URL or string instantly.
- Percent-encoding in 60 seconds: the %XX rule
- Percent-encoding chart: the characters you'll actually meet
- Which characters must be encoded — and which never are
- URL encode vs decode: two directions, one system
- Encoding in JavaScript and Python (and by hand)
- Where broken encoding bites: query strings, file names, and redirects
- Non-English characters: UTF-8 turns one letter into several %XX groups
Percent-encoding in 60 seconds: the %XX rule
URLs allow only a limited character set, so everything else gets encoded as %XX — a percent sign plus two hexadecimal digits giving the character's code. The recipe: Step 1: take the character's ASCII (or UTF-8) code. Step 2: write that code in hexadecimal, two digits. Step 3: stick a % in front. Space is ASCII 32, which is 20 in hex → %20. Ampersand is ASCII 38, hex 26 → %26. The literal percent sign itself is ASCII 37, hex 25 → %25 (yes, encoding introduces percent signs that then need encoding — that's why you encode once, at the end). Decoding is the reverse: read %XX, convert the hex to a number, look up the character. %41 → hex 41 = decimal 65 → 'A'. The whole system is deterministic and universal — every browser, server, and programming language agrees on these mappings, which is why a URL encoded on a Mac works on a Windows server in Tokyo. When you need it done without thinking, an encoder tool handles the conversion in one paste — but knowing the %XX rule lets you read encoded URLs at a glance, which is a genuine debugging superpower.
Percent-encoding chart: the characters you'll actually meet
You don't need the full 256-entry table — a dozen characters cover nearly every real-world case. Memorize these and you'll read encoded URLs fluently:
space → %20 ! → %21 " → %22 # → %23 $ → %24 % → %25
& → %26 ' → %27 ( → %28 ) → %29 * → %2A + → %2B
, → %2C / → %2F : → %3A ; → %3B = → %3D ? → %3F
@ → %40 [ → %5B ] → %5D
Notice the pattern: %20-%2F covers punctuation in the 32-47 ASCII range, %3A-%40 covers 58-64, %5B-%5D covers 91-93 — the chart is just ASCII order wearing a percent sign. The three you'll meet daily: %20 (space, in every file name with a gap in it), %26 (ampersand, whenever one appears inside a query value), and %3F (question mark inside data rather than starting the query). Case doesn't matter — %2f and %2F are identical — but uppercase is the convention and what most tools emit. Bookmark this chart or keep the encoder open in a tab; after a month of glancing at encoded URLs you'll stop needing either.
Which characters must be encoded — and which never are
Not every special character needs encoding — the rules divide characters into three camps. Never encode (unreserved): A-Z, a-z, 0-9, hyphen (-), underscore (_), period (.), tilde (~). These are always safe, everywhere in a URL. Encode when used as data (reserved): : / ? # [ ] @ ! $ & ' ( ) * + , ; = — these have structural jobs in URLs (: separates scheme, / separates path segments, ? starts the query, # starts the fragment, & separates parameters), so when they appear as data rather than structure, they must be encoded or the URL gets misparsed. Always encode: spaces, control characters, non-ASCII characters, and the characters < > " { } | \ ^ ` — these are simply illegal in URLs. The classic bug: a search box submits 'fish & chips' and the & splits it into two parameters — 'q=fish ' and ' chips=' — silently corrupting the query. Encoding the & as %26 keeps it as data. The mirror-image bug is over-encoding: encoding the : and / in 'https://example.com' produces 'https%3A%2F%2Fexample.com', a broken URL — structure stays raw, data gets encoded. That single distinction (structure raw, data encoded) resolves 90% of encoding confusion.
URL encode vs decode: two directions, one system
Encoding and decoding are inverse operations, and mixing up the direction is a perennial bug source. Encode (text → %XX) when you're putting data into a URL: a search term into a query string, a file name into a path, user input into a link you're generating. Decode (%XX → text) when you're reading data out of a URL: parsing the query string on arrival, displaying a file name from a path, logging what was actually requested. The telltale sign of a direction bug is double-encoding: 'a%2520b' — the % of %20 got encoded again into %25, producing %2520. It happens when data passes through two encoding steps (a framework encodes automatically, then your code encodes again). The fix is to encode exactly once, at the boundary where data enters the URL. The reverse bug, double-decoding, is a security issue: %2520 decoded twice becomes a space, which can smuggle characters past validation that checked the single-decoded form — validators must check the final decoded value. Rule of thumb: encode on the way out (building URLs), decode on the way in (parsing them), exactly once each. For quick checks in either direction, the encoder/decoder tool does both.
Encoding in JavaScript and Python (and by hand)
Every language ships encoding functions — use them instead of hand-rolling. JavaScript: encodeURIComponent('fish & chips') → 'fish%20%26%20chips' — this is the one you want for query values, because it encodes everything except the unreserved set. Its sibling encodeURI() is for whole URLs: it leaves structural characters (:/?#@&=+$.,;) raw and only encodes the illegal ones — using encodeURI on a query value leaves the & unencoded and breaks the query, the most common JS encoding bug. Decode with decodeURIComponent(). Python: urllib.parse.quote('fish & chips') → 'fish%20%26%20chips'; quote_plus() encodes spaces as + instead (the form-encoding convention). Unquote with urllib.parse.unquote(). By hand: for a single character, look up its ASCII code, convert to hex, prefix % — 'é' is the fun one: it's not ASCII at all, so UTF-8 encodes it as two bytes, C3 A9, giving %C3%A9 (more on that below). Whatever the language, never build the %XX strings with string replacement — one missed character class and you have a bug that only appears with certain inputs. The stdlib functions encode exactly the right set, every time.
Where broken encoding bites: query strings, file names, and redirects
Encoding bugs don't announce themselves — they corrupt silently. Query strings: 'search?q=rock & roll' splits at the & into q='rock ' plus a garbage parameter; the fix is encoding the value to 'rock%20%26%20roll'. Related: the + trap — in query strings, a literal + means space (form-encoding convention), so searching for 'C++' must encode the pluses as %2B or the server receives 'C '. File names: 'my report (final).pdf' in a URL becomes 'my%20report%20(final).pdf' — parentheses are legal unencoded in practice, but spaces must go; servers that receive raw spaces may 404 or, worse, serve the wrong file. Redirects: a redirect target containing a ? or # that isn't encoded sends the browser to a different URL than intended — open-redirect vulnerabilities live here. Fragments: everything after # never reaches the server at all — encoding a # as %23 keeps it as data, leaving it raw starts a fragment. The debugging workflow when a URL misbehaves: decode it fully with a decoder, eyeball the structure vs data split, find the character that's playing a structural role it shouldn't, and encode that one. Nine times out of ten it's an & or a #.
Non-English characters: UTF-8 turns one letter into several %XX groups
ASCII covers English; the rest of the world's text takes a two-step trip. Step 1: encode the character as UTF-8 bytes. Step 2: percent-encode each byte. So 'é' (U+00E9) becomes UTF-8 bytes C3 A9, then %C3%A9. '中' (U+4E2D) is three UTF-8 bytes E4 B8 AD → %E4%B8%AD. An emoji like 😀 is four bytes → four %XX groups. This is why encoded non-English URLs look so long: one visible character can be 9-12 encoded characters. Three things follow. First, never encode with the wrong charset — legacy pages encoded in Latin-1 produce %E9 for é instead of %C3%A9, and modern servers decode it as mojibake; UTF-8 is the standard, full stop. Second, browsers display decoded URLs in the address bar (showing 'café' while sending 'caf%C3%A9'), which is convenient until you're debugging — copy the URL from view-source or dev tools to see what's actually sent. Third, this connects to everyday publishing: URL slugs transliterate é→e precisely to avoid this byte-explosion, and QR codes containing URLs get physically denser as encoding lengthens the payload. For the conversion itself, the encoder handles UTF-8 correctly — hand-encoding multi-byte characters is an exercise, not a workflow.