Java正则表达式:排除特定标点及无空格后缀点的匹配需求
Got it, let's work through this regex problem for Java. Your goal is to strip all punctuation from a string except for /, #, $, and dots that don’t have a space right after them (like the ones in .net or asp.net). Here's how to adjust your original regex to cover that edge case with dots:
First, let’s break down what we need to remove:
- Any punctuation that isn’t
/,#,$, or a dot - Dots that are immediately followed by a space or are at the end of the string
Here’s the Java code that does exactly this:
String input = ".net, asp.net hello. world, c#"; String cleanedString = input.replaceAll("(?:(?!\\/|#|\\$|\\.)\\p{P})|(?:\\.(?=\\s|$))", ""); System.out.println(cleanedString); // Output: .net asp.net hello world c#
Let’s break down the regex to make it clear:
(?:(?!\\/|#|\\$|\\.)\\p{P}): This matches any punctuation character (\\p{P}) that isn’t/,#,$, or a dot. The(?!...)negative lookahead ensures we skip those specific excluded characters.|: Acts as an OR to combine the two patterns we want to target.(?:\\.(?=\\s|$)): This targets dots that are followed by a space (\\s) or the end of the string ($). The(?=...)positive lookahead checks what comes right after the dot to confirm it’s a space or the end of the line.
A quick note for Java specifics: We have to escape backslashes, so regex symbols like \p{P} become \\p{P}, / becomes \\/, and $ becomes \\$. The (?:...) non-capturing groups just help organize the regex without creating unnecessary capture groups that we don’t need for replacement.
内容的提问来源于stack exchange,提问作者thomas

