You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Stanza按指定语言移除TXT文件中的停用词?

多语言TXT文件停用词移除问题

我有多个不同语言的TXT文件,用下面的代码让用户指定文本语言:

lng = input("In what language is the text typed? ('ca' for catalan, 'es' for spanish, 'en' for english...)
")

我需要移除所有停用词,把处理后的文本存到新的TXT文件里。后续要做情感分析,本来想选Stanza,但不知道怎么实现停用词移除;试了速度更快的Spacy,明明已经装了es_core_web_sm模型,却报这个错:

OSError: [E050] Can't find model 'es_core_web_sm'. It doesn't seem to be a Python package or a valid path to a data directory.

附上了文件示例和模型安装情况的截图,我尝试的Spacy代码如下:

import spacy

sp = spacy.load(str(lng) + '_core_web_sm') # the inputted language is stored in 'lng'
all_stopwords = sp.Defaults.stop_words
y = open('NODUP_FILTERED_' + filename, 'r', encoding='utf-8')
txt = y.read()
for line in range(rn):
    for word in txt:
        if word in all_stopwords:
            word = ''
print(txt)

问题解决方案

一、Spacy模型加载报错+代码修正

  1. 模型问题排查
    • 先确认模型真的装好了:打开终端输pip list,看看有没有es_core_web_sm的条目。如果没有,重新装:
      • 用官方命令安装(注意匹配当前Python环境,别搞混虚拟环境):python -m spacy download es_core_web_sm
      • 要是已经下载了模型文件,直接用本地路径加载:spacy.load("/你存放模型的文件夹路径/es_core_web_sm")
  2. 代码逻辑修正
    原代码是遍历单个字符,而且直接改word=''根本不会修改原文本,正确的做法是先分词再过滤:
    import spacy
    
    # 先确保模型安装正确再加载
    nlp = spacy.load(f"{lng}_core_web_sm")
    all_stopwords = nlp.Defaults.stop_words
    
    # 用with语句安全读写文件
    with open(f'NODUP_FILTERED_{filename}', 'r', encoding='utf-8') as f:
        txt = f.read()
    
    # 分词后过滤停用词和空字符
    doc = nlp(txt)
    filtered_tokens = [token.text for token in doc if not token.is_stop and token.text.strip()]
    filtered_text = ' '.join(filtered_tokens)
    
    # 保存处理后的文本
    with open(f'FILTERED_{filename}', 'w', encoding='utf-8') as f:
        f.write(filtered_text)
    

二、Stanza实现停用词移除

Stanza处理停用词很直观,加载对应语言模型后,直接用token.is_stop判断即可:

import stanza

# 首次运行会自动下载对应语言模型,之后无需重复下载
stanza.download(lng)
nlp = stanza.Pipeline(lng)

# 读取原文件
with open(f'NODUP_FILTERED_{filename}', 'r', encoding='utf-8') as f:
    txt = f.read()

# 处理文本并过滤停用词
doc = nlp(txt)
filtered_tokens = []
for sentence in doc.sentences:
    for token in sentence.tokens:
        if not token.is_stop and token.text.strip():
            filtered_tokens.append(token.text)
filtered_text = ' '.join(filtered_tokens)

# 保存结果
with open(f'FILTERED_{filename}', 'w', encoding='utf-8') as f:
    f.write(filtered_text)

内容的提问来源于stack exchange,提问作者Questioneer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 06:36:32