移除停用词后修复单词间距异常问题
问题排查与解决
你的代码出现单词无空格拼接的核心原因是输入文本没有被spaCy正确分词为独立token,导致filtered_sentence列表中只有一个完整字符串元素,' '.join()自然不会生成分隔空格。
排查步骤
添加分词调试代码:在
remove_stopwords函数中,创建doc对象后添加打印语句,查看分词结果:doc = nlp(text) print("分词结果:", [token.text for token in doc]) # 调试用如果输出是
['formserverhappen'],说明spaCy没有识别到文本中的分隔符,把整段当成了一个token。检查输入文本的空格类型:用
repr()查看原始文本的真实格式:print(repr(negative.ctstring.iloc[0])) # 查看第一条数据的原始格式若输出不是
'form server happen',而是包含全角空格('form server happen')、不可见控制字符等,spaCy默认分词器不会将其视为分隔符。
解决方案
方案1:预处理统一空格类型
在传入spaCy前,先将所有空白字符替换为普通空格:
import re def remove_stopwords(text,nlp,custom_stop_words=None,remove_small_tokens=True,min_len=2): # 新增:统一替换所有空白字符为普通空格 text = re.sub(r'\s+', ' ', text.strip()) if custom_stop_words: nlp.Defaults.stop_words |= custom_stop_words filtered_sentence = [] doc=nlp(text) for token in doc: if not token.is_stop: if remove_small_tokens: if len(token.text) > min_len: filtered_sentence.append(token.text) else: filtered_sentence.append(token.text) return ' '.join(filtered_sentence) if filtered_sentence else None
(注:把原代码的' '.join()改成' '.join()更符合常规需求,若需要两个空格可改回)
方案2:验证自定义停用词影响
确认自定义停用词集合{"mask","mandates"}没有包含你的目标单词(form/server/happen),不过从你的输入输出看,这三个词都被保留了,所以这个因素可以排除。
额外建议
- 避免直接修改spaCy的默认停用词表(
nlp.Defaults.stop_words |= custom_stop_words),每次调用函数都会重复添加,可能导致意外问题。更稳妥的方式是创建临时停用词集合:stop_words = nlp.Defaults.stop_words.copy() if custom_stop_words: stop_words.update(custom_stop_words) # 后续判断用:if token.text not in stop_words
内容的提问来源于stack exchange,提问作者malzie_31
相关产品推荐
相关产品推荐

