You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame文本过滤误匹配问题:如何精准识别脏话词汇

解决脏话过滤中的误匹配问题(避免将子串识别为完整脏话单词)

问题核心:当前代码通过子串匹配判断文本是否包含脏话,导致像assure这类包含ass子串的正常单词被误判为脏话。要解决这个问题,需要匹配完整单词而非子串,可通过正则表达式的**单词边界\b**实现。

修改后的代码实现

import pandas as pd

# 示例脏话词汇DataFrame
df_profanity_en = pd.DataFrame({
    'word': ['bad', 'offensive', 'curse', 'vulgar', 'ass']
})

# 示例待过滤文本DataFrame
df_ed = pd.DataFrame({
    'id': [1, 2, 3, 4, 5],
    'text': ['This is a clean text. Lets assure', 'There is a bad word here', 'No profanity in this one', 'Watch your language!', 'Another clean text'],
    'predicted_emotion': ['happy', 'sad', 'neutral', 'angry', 'happy']
})

# 构建带单词边界的正则表达式,确保匹配完整单词
profane_pattern = r'\b(' + '|'.join(df_profanity_en['word'].tolist()) + r')\b'

# 过滤包含完整脏话单词的文本(大小写不敏感)
df_filtered = df_ed[df_ed['text'].str.contains(profane_pattern, case=False, na=False, regex=True)]

# 提取匹配到的脏话单词(同样用单词边界确保完整匹配)
def find_profane_word(text, profane_words):
    for word in profane_words:
        if pd.Series(text).str.contains(r'\b' + word + r'\b', case=False, regex=True).iloc[0]:
            return word
    return None

df_filtered['profane_word'] = df_filtered['text'].apply(lambda x: find_profane_word(x, df_profanity_en['word'].tolist()))

# 重置索引
df_filtered.reset_index(drop=True, inplace=True)

print(df_filtered)

输出结果

id                      text predicted_emotion profane_word
0   2  There is a bad word here               sad          bad

关键说明

  • 单词边界\b:正则中的\b匹配单词的开头或结尾(即单词与非单词字符的分界,比如空格、标点、字符串首尾),确保只有当ass作为独立单词出现时才会被匹配,不会触发assure这类包含子串的正常单词。
  • 大小写不敏感:通过case=False参数保证匹配不受大小写影响(比如匹配Ass或ASS)。
  • 扩展场景处理:如果脏话词汇包含特殊正则字符(如$、.),需要先对单词进行转义,可使用re.escape(word)处理,避免正则语法错误。例如:
    import re
    profane_pattern = r'\b(' + '|'.join(re.escape(word) for word in df_profanity_en['word'].tolist()) + r')\b'
    

内容的提问来源于stack exchange,提问作者John Angelopoulos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 05:42:50