在XSS防护中为何需将&转义为&?仅转四类字符是否有风险?
& escaping (while escaping < > ' ") leave XSS vulnerabilities in HTML content/attributes? Great question—this is a common point of confusion when implementing input sanitization, so let’s break down the risks across both of your target scenarios, plus some edge cases you might not have considered.
1. Risk in tag content (<tag>userInput</tag>)
If you only escape < > ' " but leave & unescaped, attackers can use HTML entity references to bypass your sanitization and inject malicious code.
For example, suppose your code takes user input and inserts it into a <div> without escaping &:
<div>USER_INPUT</div>
An attacker could input <script>alert('XSS')</script>. Since you didn’t escape &, the browser will parse < as < and > as >, resulting in:
<div><script>alert('XSS')</script></div>
This directly executes the script—classic stored XSS. Even worse, some browsers will parse incomplete entities like < (without the trailing semicolon) as <, expanding the attack surface even further.
2. Risk in attribute values (<tag attribute="userInput"></tag>)
You mentioned that &quot; can’t close attributes, but that’s only if you’ve already escaped & to &. If you leave & unescaped, attackers can use " (which the browser parses as ") to break out of the attribute context.
Take this example:
<input type="text" value="USER_INPUT">
An attacker inputs " onfocus="alert('XSS')". Without escaping &, the browser converts " to ", turning the HTML into:
<input type="text" value="" onfocus="alert('XSS')">
Now, when the user focuses the input, the malicious script runs. This works because the unescaped & lets the entity resolve to a quote that closes the original attribute, allowing injection of an event handler.
Why & is non-negotiable for escaping
HTML entities rely on & as their starting character. By leaving it unescaped, you’re essentially letting attackers "reconstruct" the characters you tried to block using entity encoding. Even if you escape < > ' ", the browser will happily parse entity references into those exact characters if & is left as-is.
Final takeaway
Skipping & escaping is a critical mistake. You must escape & to & along with < > ' " (and ideally other context-specific characters) to fully mitigate XSS risks in HTML contexts.
内容的提问来源于stack exchange,提问作者z3tt4

