Python正则匹配含特殊字符单词的问题及修正方法问询
Hey there! Let's sort out that regex problem you're facing. Your current pattern \w{1,15} only matches alphanumeric characters and underscores—so it's no wonder it's missing words like single-word and b**m, and correctly ignoring the !? junk (which is what we want, but we need to capture the right words too).
The Solution Regex
The pattern you need is r'[\w*-]+'. Let's break down what each part does:
[\w*-]: This character set matches any alphanumeric character (\wcovers letters, numbers, and underscores), hyphens (-), and asterisks (*). Placing the hyphen at the end of the set keeps it from being misinterpreted as a range marker (likea-z).+: This quantifier tells the regex to match one or more of the characters in the set, so it grabs entire words instead of single characters.
Updated Code
Here's your revised code that will produce the exact output you want:
import re str1 = "These should be counted as a single-word, b**m !?" match_pattern = re.findall(r'[\w*-]+', str1) print(match_pattern)
When you run this, you'll get the desired result:
['These', 'should', 'be', 'counted', 'as', 'a', 'single-word', 'b**m']
If you ever need to include other special characters in your "words" (like @ or #), just add them to the character set—for example, r'[\w*-@#]+' would include those symbols too.
内容的提问来源于stack exchange,提问作者James Rudolf

