如何用正则表达式按空格及除'-’外的标点分割文件单词?
Hey there! Let's tackle your regex problem. It sounds like you want to split text into words, where separators are any amount of whitespace plus all punctuation except -, ', and ’—and right now periods aren't being split correctly. Here's how to fix this:
The Better Approach: Match Words Instead of Splitting Separators
Instead of trying to define all the separators (which can get messy), it's often easier to directly match the words you want. This avoids edge cases like split() creating empty strings, and it's more readable.
Regex for Matching Valid Words
Use this regex to capture all sequences of characters that count as part of a word:
[\p{L}\p{N}’'-]+
Breakdown:
\p{L}: Matches any Unicode letter (works for English and other languages)\p{N}: Matches any Unicode digit’'-: Includes the allowed punctuation (hyphen, apostrophe, curly apostrophe)+: Matches one or more of these characters in a row
Example Code (Python)
import re sample_text = "Hi there! This is a test-string. Don’t forget: periods should split words here." valid_words = re.findall(r"[\p{L}\p{N}’'-]+", sample_text, re.UNICODE) print(valid_words)
This will output:
['Hi', 'there', 'This', 'is', 'a', 'test-string', 'Don’t', 'forget', 'periods', 'should', 'split', 'words', 'here']
Perfect—periods are now correctly ignored as part of words, and the split happens exactly where you want.
If You Must Use split()
If you prefer splitting instead of matching, you can define a regex that matches all your separators:
[\s\p{P}&&[^’\'-]]+
Breakdown:
\s: Matches any whitespace character (spaces, tabs, newlines)\p{P}: Matches any Unicode punctuation&&[^’\'-]: Excludes the three punctuation marks you want to keep in words+: Matches one or more of these separators in a row
Example with split()
import re sample_text = "Hello, world! Let’s try splitting with periods. It works now." split_words = re.split(r"[\s\p{P}&&[^’\'-]]+", sample_text, re.UNICODE) # Filter out any empty strings that might pop up from leading/trailing separators split_words = [word for word in split_words if word] print(split_words)
This gives the same result as the findall method, but you have to handle empty strings manually.
Key Notes
- If you only need to support English, you can replace
\p{L}witha-zA-Zand\p{N}with0-9for a simpler regex. - In Python 3, the
re.UNICODEflag is default, so you can omit it if you're using a recent version.
内容的提问来源于stack exchange,提问作者likwidmonster

