Grails 3 XSS过滤正则优化:解决表单注入过滤误匹配问题
Great question! The core issue with your current regex is that it's matching any tag that contains the substring "form" (instead of only actual <form> tags) and capturing everything between those tags—even harmless content like your refactoring paragraph where "form" appears in regular text. Let's build a precise regex that targets only malicious form injections while leaving legitimate content untouched.
Key Requirements for the Improved Regex
We need to:
- Match only actual
<form>opening tags (case-insensitive, since HTML doesn't care about tag case) - Account for optional attributes inside the opening form tag (like
class,id, or maliciousonloadhandlers) - Match the corresponding closing
</form>tag (again, case-insensitive, and allow for accidental extra characters like</form >) - Avoid matching regular text that contains "form" (like "refactoring") or other tags that happen to include "form" in their name/attributes
The Improved Regex
Here's a regex that meets these criteria, with explanations:
/<form\b[^>]*>([\s\S]*?)<\/form\b[^>]*>/gi
Breakdown of the Regex
Let's break down each part to understand how it works:
<form\b: Matches the literal<formfollowed by a word boundary (\b), ensuring we don't match tags like<transform>or<formdata>(the word boundary stops the match at the end of "form").[^>]*: Matches any character except>(non-greedy), allowing for any attributes inside the opening form tag (e.g.,<form id="malicious" onsubmit="stealData()">).>: Closes the opening tag match.([\s\S]*?): Captures all content between the opening and closing form tags.[\s\S]matches any character (including newlines), and*?makes it non-greedy so it stops at the first closing</form>tag (instead of matching all the way to the last one).<\/form\b: Matches the literal</formfollowed by a word boundary, ensuring we only match actual closing form tags (not tags like</performance>).[^>]*>: Allows for any extra characters inside the closing tag (though this is rare, it handles cases like malformed closing tags like</form >)./gi: Thegflag matches all occurrences, and theiflag makes the match case-insensitive (so it catches<FORM>,<Form>, etc.).
Testing Against Your Harmless Content
This regex will not match your refactoring paragraph because:
- There are no
<form>or</form>tags in the content. - The word "refactoring" appears inside a
<p>tag's text content, not within a tag name/attributes where the regex looks for "form".
Additional Notes for XSS Protection
While this regex fixes the over-matching issue, keep in mind:
- Regex is not perfect for HTML parsing: For more robust XSS protection, consider using a dedicated HTML sanitization library (like OWASP Java HTML Sanitizer, which integrates well with Grails) instead of relying solely on regex. Regex can miss edge cases like nested tags or malformed HTML.
- Validate and sanitize on the server: Always sanitize user input on the server side—client-side sanitization can be bypassed.
- Content Security Policy (CSP): Implement a CSP to further mitigate XSS risks by restricting which resources can load on your page.
内容的提问来源于stack exchange,提问作者Alin Pandichi

