如何使用Stanza按指定语言移除TXT文件中的停用词?
多语言TXT文件停用词移除问题
我有多个不同语言的TXT文件,用下面的代码让用户指定文本语言:
lng = input("In what language is the text typed? ('ca' for catalan, 'es' for spanish, 'en' for english...) ")
我需要移除所有停用词,把处理后的文本存到新的TXT文件里。后续要做情感分析,本来想选Stanza,但不知道怎么实现停用词移除;试了速度更快的Spacy,明明已经装了es_core_web_sm模型,却报这个错:
OSError: [E050] Can't find model 'es_core_web_sm'. It doesn't seem to be a Python package or a valid path to a data directory.
附上了文件示例和模型安装情况的截图,我尝试的Spacy代码如下:
import spacy sp = spacy.load(str(lng) + '_core_web_sm') # the inputted language is stored in 'lng' all_stopwords = sp.Defaults.stop_words y = open('NODUP_FILTERED_' + filename, 'r', encoding='utf-8') txt = y.read() for line in range(rn): for word in txt: if word in all_stopwords: word = '' print(txt)
问题解决方案
一、Spacy模型加载报错+代码修正
- 模型问题排查
- 先确认模型真的装好了:打开终端输
pip list,看看有没有es_core_web_sm的条目。如果没有,重新装:- 用官方命令安装(注意匹配当前Python环境,别搞混虚拟环境):
python -m spacy download es_core_web_sm - 要是已经下载了模型文件,直接用本地路径加载:
spacy.load("/你存放模型的文件夹路径/es_core_web_sm")
- 用官方命令安装(注意匹配当前Python环境,别搞混虚拟环境):
- 先确认模型真的装好了:打开终端输
- 代码逻辑修正
原代码是遍历单个字符,而且直接改word=''根本不会修改原文本,正确的做法是先分词再过滤:import spacy # 先确保模型安装正确再加载 nlp = spacy.load(f"{lng}_core_web_sm") all_stopwords = nlp.Defaults.stop_words # 用with语句安全读写文件 with open(f'NODUP_FILTERED_{filename}', 'r', encoding='utf-8') as f: txt = f.read() # 分词后过滤停用词和空字符 doc = nlp(txt) filtered_tokens = [token.text for token in doc if not token.is_stop and token.text.strip()] filtered_text = ' '.join(filtered_tokens) # 保存处理后的文本 with open(f'FILTERED_{filename}', 'w', encoding='utf-8') as f: f.write(filtered_text)
二、Stanza实现停用词移除
Stanza处理停用词很直观,加载对应语言模型后,直接用token.is_stop判断即可:
import stanza # 首次运行会自动下载对应语言模型,之后无需重复下载 stanza.download(lng) nlp = stanza.Pipeline(lng) # 读取原文件 with open(f'NODUP_FILTERED_{filename}', 'r', encoding='utf-8') as f: txt = f.read() # 处理文本并过滤停用词 doc = nlp(txt) filtered_tokens = [] for sentence in doc.sentences: for token in sentence.tokens: if not token.is_stop and token.text.strip(): filtered_tokens.append(token.text) filtered_text = ' '.join(filtered_tokens) # 保存结果 with open(f'FILTERED_{filename}', 'w', encoding='utf-8') as f: f.write(filtered_text)
内容的提问来源于stack exchange,提问作者Questioneer
相关产品推荐
相关产品推荐

