如何用正则re按点分割字符串?多符号分割空串原因排查
Hey there! Let's tackle your regex questions step by step.
re First off, remember that the English period (.) is a special character in regex—it matches any character except a newline. So to split on literal periods, you need to escape it or use a character class:
Escape the period with a backslash: Use
\to tell regex you mean the literal dot.import re sample_text = "Hi there.This is a sentence.Another one here." split_result = re.split(r'\.', sample_text) # Output: ['Hi there', 'This is a sentence', 'Another one here', ''](Note: The final empty string is because the original text ends with a period—we'll cover how to handle that later!)
Use a character class: Putting the period inside
[]makes it a literal character, no escape needed.split_result = re.split(r'[.]', sample_text) # Same output as above
Great question—this happens for a few common reasons:
- Consecutive delimiters: If your string has back-to-back symbols (like
Hello..WorldorHey!,,There), regex splits between each symbol, leaving an empty string in the gap. - Delimiters at the start/end: If your string starts or ends with a delimiter (like
...HelloorTest.), the split will create an empty string at the beginning or end.
How to fix it
You have two solid options:
Option 1: Filter out empty strings after splitting
Just use a list comprehension to keep only non-empty entries:
text_with_multiple_symbols = "Hello..World!,,Test,,," split_raw = re.split(r'[.!,]', text_with_multiple_symbols) clean_result = [s for s in split_raw if s] # Output: ['Hello', 'World', 'Test']
Option 2: Adjust your regex to match consecutive delimiters
Use the + quantifier to match one or more consecutive delimiters—this treats them as a single split point, so no empty strings in between:
split_with_plus = re.split(r'[.!]+', text_with_multiple_symbols) clean_result = [s for s in split_with_plus if s] # Same clean output as above
Bonus: Use re.findall instead (even simpler!)
If you only care about the non-delimiter parts of the string, skip splitting entirely and use re.findall to match all sequences of characters not in your delimiter set:
result = re.findall(r'[^.!]+', text_with_multiple_symbols) # Output: ['Hello', 'World', 'Test']
This automatically ignores empty strings since it only matches actual content.
内容的提问来源于stack exchange,提问作者Markus

