如何向stopwords添加标点并解决与词汇绑定的标点无法过滤的问题
现有代码缺陷
- 直接按空格拆分句子,与单词绑定的标点会成为单词的一部分,无法被单独识别移除
- 停用词列表同时收录大小写变体,维护成本高,判断逻辑冗余
优化后代码
# 停用词统一用小写,改为集合类型提升查询效率 my_stopwords = {'is', 'it', 'the', 'if'} # 定义需要移除的标点,可根据需求自行增删 remove_punc = '.,?!;:\'"' def prep_text(sentence): # 第一步:批量移除所有目标标点 sentence_clean = sentence.translate(str.maketrans('', '', remove_punc)) # 第二步:按空格拆分单词 words = sentence_clean.split(" ") # 第三步:过滤停用词,统一转小写判断,保留原词大小写输出 words_filtered= [word for word in words if word.lower() not in my_stopwords] return " ".join(words_filtered)
效果验证
调用prep_text('how was the game?'),输出结果为how was game,完全符合预期。
如果需要自动移除所有非文字类符号,无需手动定义标点,可以用正则方案,代码修改如下:
import re my_stopwords = {'is', 'it', 'the', 'if'} def prep_text(sentence): # 正则匹配所有非字母、非空格字符,直接删除 sentence_clean = re.sub(r'[^a-zA-Z\s]', '', sentence) words = sentence_clean.split(" ") words_filtered= [word for word in words if word.lower() not in my_stopwords] return " ".join(words_filtered)
内容的提问来源于stack exchange,提问作者Abhi Khanna
相关产品推荐
相关产品推荐

