Python遍历字符串列表移除违禁词失败,求正确实现方案
问题:移除字符串列表中的指定违禁词
我有一个食材字符串列表:
dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"]
需要移除的违禁词列表包含:
bannedWord = ['grated', 'zested', 'thinly', 'chopped', ',']
期望得到的清理后列表是:
cleaner_list = ["lemons", "cheddar cheese", "carrots"]
第一次尝试代码及结果
import re dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"] cleaner_list = [] def RemoveBannedWords(ing): pattern = re.compile("\\b(grated|zested|thinly|chopped)\\W", re.I) return pattern.sub("", ing) for ing in dirtylist: cleaner_ing = RemoveBannedWords(ing) cleaner_list.append(cleaner_ing) print(cleaner_list)
返回结果:
['lemons zested', 'cheddar cheese', 'carrots, chopped']
第二次尝试代码及结果
import re dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"] cleaner_list = [] bannedWord = ['grated', 'zested', 'thinly', 'chopped'] re_banned_words = re.compile(r"\b(" + "|".join(bannedWord) + ")\\W", re.I) def remove_words(ing): global re_banned_words return re_banned_words.sub("", ing) for ing in dirtylist: cleaner_ing = remove_words(ing) cleaner_list.append(cleaner_ing) print(cleaner_list)
返回结果:
['lemons zested', 'cheddar cheese', 'carrots, chopped']
问题分析
你的正则表达式存在以下几个问题:
- 正则中的
\W要求违禁词后面必须跟一个非单词字符,但像zested在lemons zested里是最后一个词,后面没有非单词字符,导致匹配失败无法替换。 - 两次尝试都未将违禁词列表中的逗号
,加入匹配逻辑,所以逗号无法被移除。 - 替换后会残留多余空格,比如去掉
grated后字符串开头会有空格,影响最终结果格式。
修正方案
调整正则匹配逻辑,覆盖违禁词在字符串开头、中间、结尾的情况,同时处理逗号,最后清理多余空格:
import re dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"] bannedWord = ['grated', 'zested', 'thinly', 'chopped', ','] # 构建正则:转义违禁词避免特殊字符干扰,匹配词边界或逗号 pattern = re.compile(r'\b(' + '|'.join(re.escape(word) for word in bannedWord) + r')\b|,', re.I) cleaner_list = [] for ing in dirtylist: # 替换所有违禁词 cleaned = pattern.sub('', ing) # 清理多余空格:多空格转单空格,去除首尾空格 cleaned = re.sub(r'\s+', ' ', cleaned).strip() cleaner_list.append(cleaned) print(cleaner_list)
运行结果:
['lemons', 'cheddar cheese', 'carrots']
代码说明
re.escape():处理违禁词中的特殊字符(比如逗号),避免破坏正则结构。- 正则模式:同时匹配带词边界的违禁词和独立逗号,确保所有目标内容都能被匹配替换。
- 空格清理:通过
re.sub(r'\s+', ' ', cleaned).strip()统一处理多余空格,保证结果格式整洁。
内容的提问来源于stack exchange,提问作者JimmyStrings
相关产品推荐
相关产品推荐

