Python正则表达式需保留特殊字符的改写求助
non-U.S Got it, let's sort out this regex problem for you!
The Issue with Your Original Regex
Your original regex \b(\w+)\b only matches sequences of pure alphanumeric characters (\w covers letters, numbers, and underscores). That's why terms like non-U.S get split up: the - and . are treated as word separators, breaking the term into non, U, and S—which is why your output shows non US instead of the full non-U.S.
The Solution: Adjust the Regex to Include Allowed Special Characters
We need to modify the regex to include the - and . characters that are part of your target terms, while still ensuring we're matching complete "words" (including those with internal special characters).
Here's the updated regex and code:
import re text = 'this\r\n\r\nis a non-U.S disclosures\r\n\r\n\r\r analysis agreements disclaimer.\r\n\r\n\r\nPlease keep it confidential' cleaned_text = ' '.join(re.findall(r'\b[\w.-]+\b', text)) print(cleaned_text)
What This Does
[\w.-]+: Matches one or more characters that are either alphanumeric (\w), a dot (.), or a hyphen (-). This captures terms likenon-U.Sas a single unit.\b: Ensures we're matching full word boundaries, so we don't accidentally capture partial pieces of text.' '.join(...): Takes all the matched terms and joins them with a single space, automatically replacing all the messy line breaks and extra spaces in the original text.
Expected Output
Running this code will give you exactly what you want:
this is a non-U.S disclosures analysis agreements disclaimer. Please keep it confidential
内容的提问来源于stack exchange,提问作者kabilan karunakaran

