基于Spacy实现多语言文本停用词移除的技术求助
多语言文本停用词移除方案
问题描述
我有一个包含多种语言的字符串数组,需要移除其中的停用词。示例字符串如下:
["mai fostul președinte egiptean mohamed morsi ", "em bon jovi lançou o álbum have a nice day a ", " otok škulj är en ö i kroatien den ligger i län"...]
计划支持的语言列表:
['French', 'Spanish', 'Thai', 'Russian', 'Persian', 'Indonesian', 'Arabic', 'Pushto', 'Kannada', 'Danish', 'Japanese', 'Malayalam', 'Latin', 'Romanian', 'Swedish', 'Portugese', 'English', 'Turkish', 'Tamil', 'Urdu', 'Korean', 'German', 'Greek', 'Italian', 'Chinese', 'Dutch', 'Estonian', 'Hindi']
目前使用Spacy库,尝试的代码如下:
import pandas as pd import nltk nltk.download('punkt') import spacy nlp = spacy.load("xx_ent_wiki_sm") from spacy.tokenizer import Tokenizer from nltk.corpus import stopwords from nltk.tokenize import word_tokenize doc = nlp("This is a sentence about Facebook.") print([(ent.text, ent.label) for ent in doc.ents]) all_stopwords = nlp.Defaults.stop_words all_stopwords = nlp.Defaults.stop_words data_text=df1['Text'] #here where i store my strings for x in data_text: text_tokens = word_tokenize(x) tokens_without_sw=[word for word in text_tokens if not word inall_stopwords] print(tokens_without_sw)
解决方案
现有代码的问题
xx_ent_wiki_sm是Spacy的多语言命名实体识别模型,它的停用词库非常有限,仅包含少量通用停用词,完全覆盖不了你需要的所有语言。- 混用NLTK的
word_tokenize和Spacy的停用词库,两者分词逻辑不一样,很容易出现停用词匹配不准的情况。
方案一:基于Spacy单语言模型的多语言处理
Spacy针对大部分主流语言都有官方模型,每个模型自带对应语言的精准停用词库。可以先检测文本的语言,再加载对应模型处理:
步骤1:安装所需模型
根据你要支持的语言,安装对应的Spacy模型,比如:
# 英文模型 python -m spacy download en_core_web_sm # 西班牙语模型 python -m spacy download es_core_news_sm # 法语模型 python -m spacy download fr_core_news_sm # 其他语言同理,按需安装官方或社区提供的模型
步骤2:实现语言检测+停用词移除
用langdetect库检测文本语言,再调用对应Spacy模型处理:
import pandas as pd import spacy from langdetect import detect, LangDetectException # 映射语言名称到Spacy模型名,按需补充 lang_model_map = { 'French': 'fr_core_news_sm', 'Spanish': 'es_core_news_sm', 'English': 'en_core_web_sm', 'German': 'de_core_news_sm', 'Italian': 'it_core_news_sm', 'Dutch': 'nl_core_news_sm', 'Romanian': 'ro_core_news_sm', 'Swedish': 'sv_core_news_sm', 'Danish': 'da_core_news_sm', 'Turkish': 'tr_core_news_sm', # 小语种如果有社区模型也可以加进来 } # 提前加载所有需要的模型,避免重复加载浪费资源 nlp_models = {lang: spacy.load(model) for lang, model in lang_model_map.items()} def remove_stopwords(text): try: # 检测文本语言 lang_code = detect(text) # 把langdetect返回的代码映射成语言名称 lang_code_map = { 'en':'English', 'fr':'French', 'es':'Spanish', 'de':'German', 'it':'Italian', 'nl':'Dutch', 'ro':'Romanian', 'sv':'Swedish', 'da':'Danish', 'tr':'Turkish' } lang_name = lang_code_map.get(lang_code) # 不支持的语言直接返回原文本 if not lang_name or lang_name not in nlp_models: return text # 加载对应模型处理文本 nlp = nlp_models[lang_name] doc = nlp(text) # 过滤停用词和标点 return ' '.join([token.text for token in doc if not token.is_stop and not token.is_punct]) except LangDetectException: # 无法检测语言时返回原文本 return text # 批量处理DataFrame中的文本 df1['Text_without_stopwords'] = df1['Text'].apply(remove_stopwords)
方案二:结合NLTK多语言停用词库
NLTK提供了多种语言的停用词库,适合处理更多小语种,同样配合语言检测使用:
import pandas as pd import nltk from nltk.corpus import stopwords from nltk.tokenize import word_tokenize from langdetect import detect, LangDetectException # 下载NLTK的停用词和分词资源 nltk.download('stopwords') nltk.download('punkt') # 映射语言名称到NLTK停用词库的代码 lang_stopwords_map = { 'French': 'french', 'Spanish': 'spanish', 'English': 'english', 'German': 'german', 'Italian': 'italian', 'Dutch': 'dutch', 'Swedish': 'swedish', 'Danish': 'danish', 'Turkish': 'turkish', 'Arabic': 'arabic', 'Hindi': 'hindi', # 其他语言可以参考nltk.corpus.stopwords.fileids()的输出补充 } # 提前加载所有停用词集合,加快处理速度 stopwords_dict = {lang: set(stopwords.words(code)) for lang, code in lang_stopwords_map.items()} def remove_stopwords_nltk(text): try: lang_code = detect(text) lang_code_map = { 'en':'English', 'fr':'French', 'es':'Spanish', 'de':'German', 'it':'Italian', 'nl':'Dutch', 'sv':'Swedish', 'da':'Danish', 'tr':'Turkish', 'ar':'Arabic', 'hi':'Hindi' } lang_name = lang_code_map.get(lang_code) if not lang_name or lang_name not in stopwords_dict: return text sw = stopwords_dict[lang_name] # 分词后过滤停用词和非字母内容 tokens = word_tokenize(text) return ' '.join([token for token in tokens if token.lower() not in sw and token.isalpha()]) except LangDetectException: return text # 批量处理文本 df1['Text_without_stopwords_nltk'] = df1['Text'].apply(remove_stopwords_nltk)
注意事项
- 对于Pushto、Kannada这类小语种,Spacy可能没有官方模型,NLTK也可能没有现成的停用词库,这时可以自己收集并定义自定义停用词列表。
- 语言检测对短文本可能不准,要是你的数据里短文本多,可以考虑手动标注部分语言,或者调整检测的置信度阈值。
- 匹配停用词时记得统一转成小写,避免大小写导致的匹配失败。
内容的提问来源于stack exchange,提问作者elfii
相关产品推荐
相关产品推荐

