You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历字符串列表移除违禁词失败,求正确实现方案

问题:移除字符串列表中的指定违禁词

我有一个食材字符串列表:

dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"]

需要移除的违禁词列表包含:

bannedWord = ['grated', 'zested', 'thinly', 'chopped', ',']

期望得到的清理后列表是:

cleaner_list = ["lemons", "cheddar cheese", "carrots"]

第一次尝试代码及结果

import re

dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"]
cleaner_list = []
    
def RemoveBannedWords(ing):
    pattern = re.compile("\\b(grated|zested|thinly|chopped)\\W", re.I)
    return pattern.sub("", ing)
    
for ing in dirtylist:
    cleaner_ing = RemoveBannedWords(ing)
    cleaner_list.append(cleaner_ing)
    
print(cleaner_list)

返回结果:

['lemons zested', 'cheddar cheese', 'carrots, chopped']

第二次尝试代码及结果

import re

dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"]
cleaner_list = []

bannedWord = ['grated', 'zested', 'thinly', 'chopped']
re_banned_words = re.compile(r"\b(" + "|".join(bannedWord) + ")\\W", re.I)

def remove_words(ing):
    global re_banned_words
    return re_banned_words.sub("", ing)

for ing in dirtylist:
    cleaner_ing = remove_words(ing)
    cleaner_list.append(cleaner_ing)
  
print(cleaner_list)

返回结果:

['lemons zested', 'cheddar cheese', 'carrots, chopped']

问题分析

你的正则表达式存在以下几个问题:

  1. 正则中的\W要求违禁词后面必须跟一个非单词字符,但像zested在lemons zested里是最后一个词,后面没有非单词字符,导致匹配失败无法替换。
  2. 两次尝试都未将违禁词列表中的逗号,加入匹配逻辑,所以逗号无法被移除。
  3. 替换后会残留多余空格,比如去掉grated后字符串开头会有空格,影响最终结果格式。

修正方案

调整正则匹配逻辑,覆盖违禁词在字符串开头、中间、结尾的情况,同时处理逗号,最后清理多余空格:

import re

dirtylist = ["lemons zested", "grated cheddar cheese", "carrots, thinly chopped"]
bannedWord = ['grated', 'zested', 'thinly', 'chopped', ',']

# 构建正则:转义违禁词避免特殊字符干扰,匹配词边界或逗号
pattern = re.compile(r'\b(' + '|'.join(re.escape(word) for word in bannedWord) + r')\b|,', re.I)

cleaner_list = []
for ing in dirtylist:
    # 替换所有违禁词
    cleaned = pattern.sub('', ing)
    # 清理多余空格:多空格转单空格,去除首尾空格
    cleaned = re.sub(r'\s+', ' ', cleaned).strip()
    cleaner_list.append(cleaned)

print(cleaner_list)

运行结果:

['lemons', 'cheddar cheese', 'carrots']

代码说明

  • re.escape():处理违禁词中的特殊字符(比如逗号),避免破坏正则结构。
  • 正则模式:同时匹配带词边界的违禁词和独立逗号,确保所有目标内容都能被匹配替换。
  • 空格清理:通过re.sub(r'\s+', ' ', cleaned).strip()统一处理多余空格,保证结果格式整洁。

内容的提问来源于stack exchange,提问作者JimmyStrings

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 10:18:18