You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Grails 3 XSS过滤正则优化:解决表单注入过滤误匹配问题

Fixing Overly Broad Regex for XSS Form Injection Stripping

Great question! The core issue with your current regex is that it's matching any tag that contains the substring "form" (instead of only actual <form> tags) and capturing everything between those tags—even harmless content like your refactoring paragraph where "form" appears in regular text. Let's build a precise regex that targets only malicious form injections while leaving legitimate content untouched.

Key Requirements for the Improved Regex

We need to:

  • Match only actual <form> opening tags (case-insensitive, since HTML doesn't care about tag case)
  • Account for optional attributes inside the opening form tag (like class, id, or malicious onload handlers)
  • Match the corresponding closing </form> tag (again, case-insensitive, and allow for accidental extra characters like </form >)
  • Avoid matching regular text that contains "form" (like "refactoring") or other tags that happen to include "form" in their name/attributes

The Improved Regex

Here's a regex that meets these criteria, with explanations:

/<form\b[^>]*>([\s\S]*?)<\/form\b[^>]*>/gi

Breakdown of the Regex

Let's break down each part to understand how it works:

  • <form\b: Matches the literal <form followed by a word boundary (\b), ensuring we don't match tags like <transform> or <formdata> (the word boundary stops the match at the end of "form").
  • [^>]*: Matches any character except > (non-greedy), allowing for any attributes inside the opening form tag (e.g., <form id="malicious" onsubmit="stealData()">).
  • >: Closes the opening tag match.
  • ([\s\S]*?): Captures all content between the opening and closing form tags. [\s\S] matches any character (including newlines), and *? makes it non-greedy so it stops at the first closing </form> tag (instead of matching all the way to the last one).
  • <\/form\b: Matches the literal </form followed by a word boundary, ensuring we only match actual closing form tags (not tags like </performance>).
  • [^>]*>: Allows for any extra characters inside the closing tag (though this is rare, it handles cases like malformed closing tags like </form >).
  • /gi: The g flag matches all occurrences, and the i flag makes the match case-insensitive (so it catches <FORM>, <Form>, etc.).

Testing Against Your Harmless Content

This regex will not match your refactoring paragraph because:

  • There are no <form> or </form> tags in the content.
  • The word "refactoring" appears inside a <p> tag's text content, not within a tag name/attributes where the regex looks for "form".

Additional Notes for XSS Protection

While this regex fixes the over-matching issue, keep in mind:

  • Regex is not perfect for HTML parsing: For more robust XSS protection, consider using a dedicated HTML sanitization library (like OWASP Java HTML Sanitizer, which integrates well with Grails) instead of relying solely on regex. Regex can miss edge cases like nested tags or malformed HTML.
  • Validate and sanitize on the server: Always sanitize user input on the server side—client-side sanitization can be bypassed.
  • Content Security Policy (CSP): Implement a CSP to further mitigate XSS risks by restricting which resources can load on your page.

内容的提问来源于stack exchange,提问作者Alin Pandichi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:13:01