You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止spaCy拆分含数字的混合字符串时误删停用词片段?

解决spaCy处理数字字母混合字符串时误删停用词片段的问题

遇到的核心问题是:spaCy的分词器会把数字和字母混合的字符串(比如co555in)拆分成多个token,其中拆分出的停用词片段(比如in)会被误删,导致输出不符合预期。官方已说明这是特性而非Bug,以下提供两种变通方案:

方案一:修改分词规则,避免数字拆分字母组合

spaCy默认的分词规则会在字母和数字之间拆分,我们可以修改infix规则,让数字和字母连在一起的部分不被拆分,这样整个混合字符串会被当成一个完整token,自然不会触发停用词过滤。

代码实现:

import spacy
from spacy.util import compile_infix_regex

# 加载英文模型(根据你的需求替换对应语言)
nlp = spacy.load("en_core_web_sm")

# 调整infix规则,移除字母与数字间的拆分逻辑
infixes = nlp.Defaults.infixes
# 过滤掉匹配字母后接数字、数字后接字母的规则
infixes = [infix for infix in infixes 
           if not infix.startswith(r"(?<=[0-9])") 
           and not infix.startswith(r"(?<=[a-zA-Z])")]
infix_re = compile_infix_regex(infixes)

# 替换tokenizer的infix匹配器
nlp.tokenizer.infix_finditer = infix_re.finditer

# 处理DataFrame的列
df[col] = df[col].apply(lambda text: 
        "".join(token.lemma_ for token in nlp(text) 
        if not token.is_stop))

这个方案从根源上避免了拆分问题,适合大部分数字字母混合的场景。

方案二:过滤停用词时增加额外判断

如果不想修改分词规则,可以在过滤停用词时,检查该停用词token是否和其他字母/数字连在一起(即不是独立的单词),如果是则保留该片段。

代码实现:

import spacy

nlp = spacy.load("en_core_web_sm")

def safe_remove_stopwords(text):
    doc = nlp(text)
    result = []
    for token in doc:
        if not token.is_stop:
            result.append(token.text)
        else:
            # 获取token在原始文本中的位置
            start_idx = token.idx
            end_idx = start_idx + len(token.text)
            # 检查token前后是否有字母/数字(判断是否为连在一起的片段)
            has_prev_char = start_idx > 0 and text[start_idx-1].isalnum()
            has_next_char = end_idx < len(text) and text[end_idx].isalnum()
            if has_prev_char or has_next_char:
                result.append(token.text)
    return "".join(result)

# 应用到DataFrame
df[col] = df[col].apply(safe_remove_stopwords)

这个方案保留了原有的分词逻辑,仅在过滤时做特殊判断,适合需要保留其他拆分场景的需求。

内容的提问来源于stack exchange,提问作者Jared

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 08:07:34