Python正则分割代码调试:移除输出列表中的空字符串
调试Python代码去除多余空字符串
原代码与问题
原代码如下:
import re text = "Hello there." word_list = [] for word in text.split(): tmp = re.split(r'(\W+)', word) word_list.extend(tmp) print(word_list)
当前输出:['Hello', 'there', '.', ''],末尾出现多余空字符串,不符合预期。
期望输出:['Hello', 'there', '.']
问题原因
使用带捕获组的re.split()时,如果正则表达式匹配到了字符串的末尾,会在拆分结果的最后生成一个空字符串。比如处理"there."时,re.split(r'(\W+)', 'there.')会返回['there', '.', '']——因为最后的.匹配在字符串末尾,后续无内容,所以追加了空串。
解决方案
方案一:过滤空字符串
在将拆分结果加入列表时,只保留非空元素,修改循环部分即可:
import re text = "Hello there." word_list = [] for word in text.split(): tmp = re.split(r'(\W+)', word) # 过滤空字符串后再添加到列表 word_list.extend(item for item in tmp if item) print(word_list)
运行后输出:['Hello', 'there', '.'],符合预期。
方案二:改用re.findall简化逻辑
如果不需要先按空白拆分的步骤,可以直接用re.findall匹配所有单词和标点,同时跳过空白字符:
import re text = "Hello there." # 匹配单词或标点,同时排除空白字符 word_list = re.findall(r'\w+|[^\w\s]', text) print(word_list)
这个正则表达式\w+|[^\w\s]表示匹配一个或多个单词字符,或者匹配非单词且非空白的字符,直接得到预期结果。
内容的提问来源于stack exchange,提问作者user13235761
相关产品推荐
相关产品推荐

