You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Spacy实现多语言文本停用词移除的技术求助

多语言文本停用词移除方案

问题描述

我有一个包含多种语言的字符串数组,需要移除其中的停用词。示例字符串如下:

["mai fostul președinte egiptean mohamed morsi ", "em bon jovi lançou o álbum have a nice day a ", " otok škulj är en ö i kroatien den ligger i län"...]

计划支持的语言列表:

['French',
 'Spanish',
 'Thai',
 'Russian',
 'Persian',
 'Indonesian',
 'Arabic',
 'Pushto',
 'Kannada',
 'Danish',
 'Japanese',
 'Malayalam',
 'Latin',
 'Romanian',
 'Swedish',
 'Portugese',
 'English',
 'Turkish',
 'Tamil',
 'Urdu',
 'Korean',
 'German',
 'Greek',
 'Italian',
 'Chinese',
 'Dutch',
 'Estonian',
 'Hindi']

目前使用Spacy库,尝试的代码如下:

import pandas as pd
import nltk
nltk.download('punkt')
import spacy
nlp = spacy.load("xx_ent_wiki_sm")
from spacy.tokenizer import Tokenizer
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize

doc = nlp("This is a sentence about Facebook.")
print([(ent.text, ent.label) for ent in doc.ents])
all_stopwords = nlp.Defaults.stop_words
all_stopwords = nlp.Defaults.stop_words

data_text=df1['Text'] #here where i store my strings 

for x in data_text:
    text_tokens = word_tokenize(x)
    tokens_without_sw=[word for word in text_tokens if not word inall_stopwords]

print(tokens_without_sw)

解决方案

现有代码的问题

  • xx_ent_wiki_sm是Spacy的多语言命名实体识别模型,它的停用词库非常有限,仅包含少量通用停用词,完全覆盖不了你需要的所有语言。
  • 混用NLTK的word_tokenize和Spacy的停用词库,两者分词逻辑不一样,很容易出现停用词匹配不准的情况。

方案一:基于Spacy单语言模型的多语言处理

Spacy针对大部分主流语言都有官方模型,每个模型自带对应语言的精准停用词库。可以先检测文本的语言,再加载对应模型处理:

步骤1:安装所需模型

根据你要支持的语言,安装对应的Spacy模型,比如:

# 英文模型
python -m spacy download en_core_web_sm
# 西班牙语模型
python -m spacy download es_core_news_sm
# 法语模型
python -m spacy download fr_core_news_sm
# 其他语言同理,按需安装官方或社区提供的模型

步骤2:实现语言检测+停用词移除

用langdetect库检测文本语言,再调用对应Spacy模型处理:

import pandas as pd
import spacy
from langdetect import detect, LangDetectException

# 映射语言名称到Spacy模型名,按需补充
lang_model_map = {
    'French': 'fr_core_news_sm',
    'Spanish': 'es_core_news_sm',
    'English': 'en_core_web_sm',
    'German': 'de_core_news_sm',
    'Italian': 'it_core_news_sm',
    'Dutch': 'nl_core_news_sm',
    'Romanian': 'ro_core_news_sm',
    'Swedish': 'sv_core_news_sm',
    'Danish': 'da_core_news_sm',
    'Turkish': 'tr_core_news_sm',
    # 小语种如果有社区模型也可以加进来
}

# 提前加载所有需要的模型,避免重复加载浪费资源
nlp_models = {lang: spacy.load(model) for lang, model in lang_model_map.items()}

def remove_stopwords(text):
    try:
        # 检测文本语言
        lang_code = detect(text)
        # 把langdetect返回的代码映射成语言名称
        lang_code_map = {
            'en':'English', 'fr':'French', 'es':'Spanish', 'de':'German', 
            'it':'Italian', 'nl':'Dutch', 'ro':'Romanian', 'sv':'Swedish', 
            'da':'Danish', 'tr':'Turkish'
        }
        lang_name = lang_code_map.get(lang_code)
        # 不支持的语言直接返回原文本
        if not lang_name or lang_name not in nlp_models:
            return text
        # 加载对应模型处理文本
        nlp = nlp_models[lang_name]
        doc = nlp(text)
        # 过滤停用词和标点
        return ' '.join([token.text for token in doc if not token.is_stop and not token.is_punct])
    except LangDetectException:
        # 无法检测语言时返回原文本
        return text

# 批量处理DataFrame中的文本
df1['Text_without_stopwords'] = df1['Text'].apply(remove_stopwords)

方案二:结合NLTK多语言停用词库

NLTK提供了多种语言的停用词库,适合处理更多小语种,同样配合语言检测使用:

import pandas as pd
import nltk
from nltk.corpus import stopwords
from nltk.tokenize import word_tokenize
from langdetect import detect, LangDetectException

# 下载NLTK的停用词和分词资源
nltk.download('stopwords')
nltk.download('punkt')

# 映射语言名称到NLTK停用词库的代码
lang_stopwords_map = {
    'French': 'french',
    'Spanish': 'spanish',
    'English': 'english',
    'German': 'german',
    'Italian': 'italian',
    'Dutch': 'dutch',
    'Swedish': 'swedish',
    'Danish': 'danish',
    'Turkish': 'turkish',
    'Arabic': 'arabic',
    'Hindi': 'hindi',
    # 其他语言可以参考nltk.corpus.stopwords.fileids()的输出补充
}

# 提前加载所有停用词集合,加快处理速度
stopwords_dict = {lang: set(stopwords.words(code)) for lang, code in lang_stopwords_map.items()}

def remove_stopwords_nltk(text):
    try:
        lang_code = detect(text)
        lang_code_map = {
            'en':'English', 'fr':'French', 'es':'Spanish', 'de':'German', 
            'it':'Italian', 'nl':'Dutch', 'sv':'Swedish', 'da':'Danish', 
            'tr':'Turkish', 'ar':'Arabic', 'hi':'Hindi'
        }
        lang_name = lang_code_map.get(lang_code)
        if not lang_name or lang_name not in stopwords_dict:
            return text
        sw = stopwords_dict[lang_name]
        # 分词后过滤停用词和非字母内容
        tokens = word_tokenize(text)
        return ' '.join([token for token in tokens if token.lower() not in sw and token.isalpha()])
    except LangDetectException:
        return text

# 批量处理文本
df1['Text_without_stopwords_nltk'] = df1['Text'].apply(remove_stopwords_nltk)

注意事项

  • 对于Pushto、Kannada这类小语种,Spacy可能没有官方模型,NLTK也可能没有现成的停用词库,这时可以自己收集并定义自定义停用词列表。
  • 语言检测对短文本可能不准,要是你的数据里短文本多,可以考虑手动标注部分语言,或者调整检测的置信度阈值。
  • 匹配停用词时记得统一转成小写,避免大小写导致的匹配失败。

内容的提问来源于stack exchange,提问作者elfii

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 01:20:35