如何用Python移除字符串指定单词列表?解决替换空格冗余问题
解决Python移除指定单词后产生多余空格的问题
你遇到的多余空格问题,本质是str.replace()只会删除目标单词本身,但单词前后的空格、标点符号会保留,导致出现连续空格或标点重复的情况(比如"RegExr Yeah was"变成"RegExr was","yippe, ow, ouch"变成"yippe, , ouch")。
解决方案:正则精确匹配+格式清理
用正则表达式匹配独立单词(避免误删非目标内容,比如不会把how里的ow删掉),再清理替换后残留的格式问题,分两步实现:
- 匹配并删除所有目标单词;
- 清理连续空格、标点前后的冗余空格。
示例代码
import re words_to_remove = ['gosh', 'no', 'oh', 'Yep', 'ow', 'well', 'goodness', 'Yeah'] test_data = """RegExr Yeah was created by gskinner.com. yippe, ow, ouch, gosh Yeah oh, goodness, oh well, oh no, how can I do wonders in this world. Yep, it is out of the world. Edit the Expression & Text, to-see matches. Roll; over$ matches% or* the expr@ession for details. PCRE & JavaScript flavors of RegEx are supported. Validate your expression with Tests mode. The side bar includes a Cheatsheet, full Reference, and Help. You can also Save & Share with the Community and view patterns you create or favorite in My Patterns. Explore results with the Tools below. Replace & List output custom results. Details lists capture groups. Explain describes your expression in plain English. """ # 构建正则模式:用\b确保匹配独立单词,re.escape处理特殊字符 pattern = r'\b(' + '|'.join(re.escape(word) for word in words_to_remove) + r')\b' # 第一步:删除目标单词 cleaned_data = re.sub(pattern, '', test_data) # 第二步:清理格式冗余 cleaned_data = re.sub(r'\s{2,}', ' ', cleaned_data) # 连续空格转单个 cleaned_data = re.sub(r'\s+([,.!?:;])', r'\1', cleaned_data) # 去掉标点前的空格 cleaned_data = re.sub(r'([,.!?:;])\s+', r'\1 ', cleaned_data) # 标点后保留单个空格 cleaned_data = '\n'.join(line.strip() for line in cleaned_data.splitlines()) # 清理行首尾空格 print(cleaned_data)
效果说明
处理后文本会保留原有换行和标点格式,无多余空格:
- 原第一行
"RegExr Yeah was created by gskinner.com."变为"RegExr was created by gskinner.com." - 原第二行冗余部分会被清理为
"yippe, ouch, how can I do wonders in this world. it is out of the world."
额外注意
- 如果需要大小写不敏感匹配(比如同时删除
yeah和Yeah),可在re.sub中添加flags=re.IGNORECASE参数; re.escape用于处理单词中的特殊字符(如.、*),确保正则匹配准确。
内容的提问来源于stack exchange,提问作者user3734568
相关产品推荐
相关产品推荐

