如何用Python移除含独立禁用词的句子而非含禁用词片段的句子
解决独立禁用词匹配与句子过滤问题
你的需求是只移除包含独立禁用词的句子,而非禁用词作为其他单词组成部分的情况,当前代码的核心问题在于直接用子串匹配(bad_word in line),这会误判像"ongoing"里的"on"这类场景。下面是具体的解决方案:
原代码的问题分析
- 禁用词处理逻辑冗余且错误:你先读取禁用词列表并做了strip处理,又循环插入小写版本,导致列表同时存在原词和小写词,可能引发重复判断。
- 匹配方式错误:子串匹配无法区分独立单词和单词中的子部分,完全不符合你的需求。
正确实现方案
我们可以用正则表达式的**单词边界(\b)**来确保匹配的是独立单词,同时通过re.IGNORECASE忽略大小写,比如"On"和"on"都会被正确识别。另外要注意用re.escape()转义禁用词中的特殊字符(比如禁用词是"don't"或者"."时,避免正则语法冲突)。
修正后的代码
import re def clean_sentences(sentences_path, bad_words_path, outfile_path, badfile_path): # 读取并处理禁用词:去除换行、过滤空行,转小写统一匹配规则 with open(bad_words_path, 'r') as f: bad_words = [word.strip().lower() for word in f.readlines() if word.strip()] # 构建正则表达式:匹配任意一个独立禁用词,忽略大小写 # 用re.escape处理特殊字符,避免正则语法错误 pattern = re.compile( r'\b(' + '|'.join(re.escape(word) for word in bad_words) + r')\b', re.IGNORECASE ) # 遍历句子文件,执行过滤逻辑 with open(sentences_path, 'r') as oldfile, \ open(outfile_path, 'w') as newfile, \ open(badfile_path, 'w') as badfile: for line in oldfile: stripped_line = line.strip() if not stripped_line: # 跳过空行,避免无效写入 continue # 检查句子中是否存在独立禁用词 if pattern.search(line): badfile.write(line) else: newfile.write(line) # 调用示例 clean_sentences('sentences.txt', 'bad_words.txt', 'outfile.txt', 'badfile.txt')
代码说明
- 禁用词处理:读取后去除空白和换行,转成小写(配合正则的忽略大小写,确保全场景匹配),同时过滤空行避免无效规则。
- 正则模式构建:用
|连接所有禁用词,并用\b包裹,确保每个禁用词都是独立单词;re.escape()处理特殊字符,避免正则语法错误。 - 句子过滤:用
pattern.search(line)检查句子中是否存在匹配的独立禁用词,存在则写入badfile,否则写入outfile。
测试你的示例
如果你的sentences.txt内容是:
Learning Python is an ongoing task I practice on and off I do it offline On weekdays i practice the most In weekends I am off
bad_words.txt内容是:
on off
运行后:
outfile.txt会保留:Learning Python is an ongoing task I do it offlinebadfile.txt会包含:I practice on and off On weekdays i practice the most In weekends I am off
完全符合你的需求:"ongoing"和"offline"里的子串不会被误判,只有独立的"on"、"off"、"On"会触发移除。
内容的提问来源于stack exchange,提问作者Jack Johnson
相关产品推荐
相关产品推荐

