You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向stopwords添加标点并解决与词汇绑定的标点无法过滤的问题

现有代码缺陷

  • 直接按空格拆分句子,与单词绑定的标点会成为单词的一部分,无法被单独识别移除
  • 停用词列表同时收录大小写变体,维护成本高,判断逻辑冗余

优化后代码

# 停用词统一用小写,改为集合类型提升查询效率
my_stopwords =  {'is', 'it', 'the', 'if'}
# 定义需要移除的标点,可根据需求自行增删
remove_punc = '.,?!;:\'"'

def prep_text(sentence):
    # 第一步:批量移除所有目标标点
    sentence_clean = sentence.translate(str.maketrans('', '', remove_punc))
    # 第二步:按空格拆分单词
    words = sentence_clean.split(" ")
    # 第三步:过滤停用词,统一转小写判断,保留原词大小写输出
    words_filtered= [word for word in words if word.lower() not in my_stopwords]
    return " ".join(words_filtered)

效果验证

调用prep_text('how was the game?'),输出结果为how was game,完全符合预期。

如果需要自动移除所有非文字类符号,无需手动定义标点,可以用正则方案,代码修改如下:

import re
my_stopwords =  {'is', 'it', 'the', 'if'}

def prep_text(sentence):
    # 正则匹配所有非字母、非空格字符,直接删除
    sentence_clean = re.sub(r'[^a-zA-Z\s]', '', sentence)
    words = sentence_clean.split(" ")
    words_filtered= [word for word in words if word.lower() not in my_stopwords]
    return " ".join(words_filtered)

内容的提问来源于stack exchange,提问作者Abhi Khanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 18:30:04