DataFrame文本过滤误匹配问题:如何精准识别脏话词汇
解决脏话过滤中的误匹配问题(避免将子串识别为完整脏话单词)
问题核心:当前代码通过子串匹配判断文本是否包含脏话,导致像assure这类包含ass子串的正常单词被误判为脏话。要解决这个问题,需要匹配完整单词而非子串,可通过正则表达式的**单词边界\b**实现。
修改后的代码实现
import pandas as pd # 示例脏话词汇DataFrame df_profanity_en = pd.DataFrame({ 'word': ['bad', 'offensive', 'curse', 'vulgar', 'ass'] }) # 示例待过滤文本DataFrame df_ed = pd.DataFrame({ 'id': [1, 2, 3, 4, 5], 'text': ['This is a clean text. Lets assure', 'There is a bad word here', 'No profanity in this one', 'Watch your language!', 'Another clean text'], 'predicted_emotion': ['happy', 'sad', 'neutral', 'angry', 'happy'] }) # 构建带单词边界的正则表达式,确保匹配完整单词 profane_pattern = r'\b(' + '|'.join(df_profanity_en['word'].tolist()) + r')\b' # 过滤包含完整脏话单词的文本(大小写不敏感) df_filtered = df_ed[df_ed['text'].str.contains(profane_pattern, case=False, na=False, regex=True)] # 提取匹配到的脏话单词(同样用单词边界确保完整匹配) def find_profane_word(text, profane_words): for word in profane_words: if pd.Series(text).str.contains(r'\b' + word + r'\b', case=False, regex=True).iloc[0]: return word return None df_filtered['profane_word'] = df_filtered['text'].apply(lambda x: find_profane_word(x, df_profanity_en['word'].tolist())) # 重置索引 df_filtered.reset_index(drop=True, inplace=True) print(df_filtered)
输出结果
id text predicted_emotion profane_word 0 2 There is a bad word here sad bad
关键说明
- 单词边界
\b:正则中的\b匹配单词的开头或结尾(即单词与非单词字符的分界,比如空格、标点、字符串首尾),确保只有当ass作为独立单词出现时才会被匹配,不会触发assure这类包含子串的正常单词。 - 大小写不敏感:通过
case=False参数保证匹配不受大小写影响(比如匹配Ass或ASS)。 - 扩展场景处理:如果脏话词汇包含特殊正则字符(如
$、.),需要先对单词进行转义,可使用re.escape(word)处理,避免正则语法错误。例如:import re profane_pattern = r'\b(' + '|'.join(re.escape(word) for word in df_profanity_en['word'].tolist()) + r')\b'
内容的提问来源于stack exchange,提问作者John Angelopoulos
相关产品推荐
相关产品推荐

