如何编写匹配SQL注释代码的Java正则表达式?
Fixing SQL Comment Removal with Java Regex
Alright, let's tackle your problem of stripping all SQL comments while ignoring comment markers that sit inside string literals. Your current regex (--.*)|(/\*(.*\n)*\*/) has a couple of critical flaws that we'll fix here.
What's Wrong with the Original Pattern?
- Greedy multi-line comment matching: The
(/\*(.*\n)*\*/)section uses greedy matching, which will latch onto the first/*and only stop at the very last*/in your SQL. This means it'll swallow valid SQL code that sits between multiple comments. - No string literal awareness: It doesn't distinguish between comment markers in actual comments vs. those inside quotes (like
SELECT 'This -- is a string' FROM table). As a result, it'll incorrectly treat parts of strings as comments.
The Solution: A Robust Regex
We need a regex that first identifies and preserves string literals, then targets and removes both single-line (--) and multi-line (/* */) comments. Here's the corrected pattern:
String regex = "('(?:\\\\.|[^'])*'|\"(?:\\\\.|[^\"])*\")|(--.*)|(/\\*[\\s\\S]*?\\*/)";
Let's break down each component:
- String literals:
('(?:\\\\.|[^'])*'|\"(?:\\\\.|[^\"])*\")- Matches both single-quoted and double-quoted strings.
(?:\\\\.|[^'])handles escaped quotes (like'It\'s a test') by either matching an escaped character (\\\\.) or any non-quote character ([^']).- This group is prioritized so we can keep strings intact instead of mistaking their content for comments.
- Single-line comments:
(--.*)- Matches everything from
--to the end of the current line.
- Matches everything from
- Multi-line comments:
(/\\*[\\s\\S]*?\\*/)[\s\S]matches every character (including newlines, which.*doesn't cover by default).- The
*?makes it non-greedy, so it stops at the first*/instead of the last one—preventing it from eating valid SQL between separate comments.
Java Code Implementation
To use this regex, we'll leverage Matcher.appendReplacement to retain string literals and discard comments:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class SQLCommentCleaner { public static void main(String[] args) { String originalSQL = "/* SELECT * FROM Table1; -- Comment in comment Another SQL code */ SQL-code -- One line comment /**/ -- Yet one comment /* And yet one comment */"; // The refined regex pattern String commentRegex = "('(?:\\\\.|[^'])*'|\"(?:\\\\.|[^\"])*\")|(--.*)|(/\\*[\\s\\S]*?\\*/)"; Pattern pattern = Pattern.compile(commentRegex); Matcher matcher = pattern.matcher(originalSQL); StringBuilder cleanedSQL = new StringBuilder(); while (matcher.find()) { // If we matched a string literal (group 1), keep it; else skip the comment if (matcher.group(1) != null) { cleanedSQL.append(matcher.group(1)); } } // Append any remaining non-matched content (valid SQL) matcher.appendTail(cleanedSQL); // Trim leftover whitespace from removed comments System.out.println("Cleaned SQL: " + cleanedSQL.toString().trim()); } }
Example Output
Running this code on your sample input will produce:
SQL-code
Notes for Edge Cases
- This regex handles most standard SQL string scenarios, but some databases have unique syntax (e.g., PostgreSQL's dollar-quoted strings). If you're working with a specific DB, you may need to adjust the string matching section.
- Always test with your actual SQL queries to ensure edge cases (like nested comment markers inside strings) are handled correctly.
内容的提问来源于stack exchange,提问作者Denis Steinman
相关产品推荐
相关产品推荐

