You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复Python停用词移除函数中的IndentationError缩进错误?

修复后的remove_stopwords函数代码
def remove_stopwords(text, nlp, custom_stop_words=None, remove_small_tokens=True, min_len=2):
    # 临时处理自定义停用词,避免修改全局默认集合
    stop_words = nlp.Defaults.stop_words.copy()
    if custom_stop_words:
        stop_words.update(custom_stop_words)
    
    filtered_sentence = []
    doc = nlp(text)
    for token in doc:
        # 跳过停用词
        if token.text.lower() not in stop_words:
            # 处理短词过滤逻辑
            if remove_small_tokens:
                if len(token.text) > min_len:
                    filtered_sentence.append(token.text)
            else:
                filtered_sentence.append(token.text)
    
    # 根据结果返回对应值
    return " ".join(filtered_sentence) if filtered_sentence else None
关键修复点说明
  • 修复缩进与语法错误:
    原代码中else None缩进错误,且if len(filtered_sentence)>0缺少冒号,同时return语句被错误放在循环内部,导致函数提前终止。修复后将return逻辑移到循环结束后,用三元表达式简洁实现判断返回逻辑。

  • 避免全局修改停用词:
    原代码直接修改nlp.Defaults.stop_words会改变spaCy模型的全局停用词集合,影响其他使用该nlp实例的代码。修复后通过复制默认停用词集合,再添加自定义停用词,仅在当前函数内生效。

  • 修正停用词判断逻辑:
    原代码依赖token.is_stop判断,但直接修改全局停用词后,后续token的is_stop才会生效,改为直接判断token文本是否在临时停用词集合中,逻辑更可靠,同时统一转为小写避免大小写匹配问题。

  • 理顺短词过滤逻辑:
    原代码中短词过滤的缩进混乱,修复后明确层级:先判断是否为停用词,再根据remove_small_tokens参数决定是否过滤短词。

内容的提问来源于stack exchange,提问作者socialscientist90

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 17:40:47