如何修复Python停用词移除函数中的IndentationError缩进错误?
修复后的
remove_stopwords函数代码 def remove_stopwords(text, nlp, custom_stop_words=None, remove_small_tokens=True, min_len=2): # 临时处理自定义停用词,避免修改全局默认集合 stop_words = nlp.Defaults.stop_words.copy() if custom_stop_words: stop_words.update(custom_stop_words) filtered_sentence = [] doc = nlp(text) for token in doc: # 跳过停用词 if token.text.lower() not in stop_words: # 处理短词过滤逻辑 if remove_small_tokens: if len(token.text) > min_len: filtered_sentence.append(token.text) else: filtered_sentence.append(token.text) # 根据结果返回对应值 return " ".join(filtered_sentence) if filtered_sentence else None
关键修复点说明
修复缩进与语法错误:
原代码中else None缩进错误,且if len(filtered_sentence)>0缺少冒号,同时return语句被错误放在循环内部,导致函数提前终止。修复后将return逻辑移到循环结束后,用三元表达式简洁实现判断返回逻辑。避免全局修改停用词:
原代码直接修改nlp.Defaults.stop_words会改变spaCy模型的全局停用词集合,影响其他使用该nlp实例的代码。修复后通过复制默认停用词集合,再添加自定义停用词,仅在当前函数内生效。修正停用词判断逻辑:
原代码依赖token.is_stop判断,但直接修改全局停用词后,后续token的is_stop才会生效,改为直接判断token文本是否在临时停用词集合中,逻辑更可靠,同时统一转为小写避免大小写匹配问题。理顺短词过滤逻辑:
原代码中短词过滤的缩进混乱,修复后明确层级:先判断是否为停用词,再根据remove_small_tokens参数决定是否过滤短词。
内容的提问来源于stack exchange,提问作者socialscientist90
相关产品推荐
相关产品推荐

