You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spacy文本分类器能否学习识别特定顺序的连续词语?

Spacy文本分类器无法区分词序:"jhon died" vs "died jhon"?

问题背景

要验证Spacy的文本分类器能否识别特定顺序的连续词组合"jhon died",但无论调整训练样本标签还是训练参数,模型始终无法只匹配"jhon died"而排除"died jhon",疑问:Spacy的textcat组件是否在分类时不考虑token顺序?

数据集详情

训练、评估及测试集均重复以下4个样本:

rows.append(["jhon died", 1])
rows.append(["died jhon", 0])
rows.append(["died", 0])
rows.append(["jhon", 0])

数据集总规模76条,拆分比例:训练集57条,开发集11条,测试集8条。

数据集生成代码

db = spacy.tokens.DocBin()
docs = []
for doc, label in nlp.pipe(data, as_tuples=True):
    doc.cats["POS"] = label == 1
    doc.cats["NEG"] = label == 0
    db.add(doc)

db.to_disk(outfile)

训练命令

python -m spacy init config  --lang en --pipeline textcat --optimize efficiency --force config.cfg

测试代码

texts = ["jhon", "jhon died", "died", "died jhon", "died fast", "fast jhon"]

nlp = spacy.load("./model/model-best")
for text in texts:
    doc = nlp(text)
    diff = doc.cats['POS'] - doc.cats['NEG']
    print("yes" if diff > 0 else ("no" if diff < 0 else "neither") ,  "-",  text, doc.cats)

测试结果

初始训练结果

no - jhon {'POS': 0.1631753146648407, 'NEG': 0.8368247151374817}
no - jhon died {'POS': 0.4730854034423828, 'NEG': 0.5269145965576172}
no - died {'POS': 0.1631753146648407, 'NEG': 0.8368247151374817}
no - died jhon {'POS': 0.4730854034423828, 'NEG': 0.5269145965576172}
no - died fast {'POS': 0.1631753146648407, 'NEG': 0.8368247151374817}
no - fast jhon {'POS': 0.1631753146648407, 'NEG': 0.8368247151374817}

调整"died jhon"标签后的结果

no - jhon {'POS': 0.21423980593681335, 'NEG': 0.785760223865509}
yes - jhon died {'POS': 0.8561566472053528, 'NEG': 0.1438433676958084}
no - died {'POS': 0.21423980593681335, 'NEG': 0.785760223865509}
yes - died jhon {'POS': 0.8561566472053528, 'NEG': 0.1438433676958084}
no - died fast {'POS': 0.21423980593681335, 'NEG': 0.785760223865509}
no - fast jhon {'POS': 0.21423980593681335, 'NEG': 0.785760223865509}

期望结果

no - jhon {...}
yes - jhon died {...}
no - died {...}
no - died jhon {...}
no - died fast {...} // Result doesn't matter here.
no - fast jhon {...} // Result doesn't matter here.

问题原因与解决方案

原因

你使用的默认textcat组件基于**词袋模型(Bag-of-Words)**架构,这类模型只统计文本中词的出现频率,完全忽略词的顺序。因此"jhon died"和"died jhon"的词袋特征完全一致,模型无法区分两者。

解决方案

要让模型捕捉词序信息,可选择以下两种方案:

  1. 使用Transformer-based文本分类器(推荐)
    重新初始化配置时,选择textcat_trf管道,Transformer模型天生具备捕捉序列语义的能力,能轻松区分词序差异:

    python -m spacy init config --lang en --pipeline textcat_trf --optimize efficiency --force config.cfg
    

    保持数据集生成和测试代码不变,重新训练即可得到符合期望的结果。

  2. 自定义添加n-gram特征
    修改textcat组件的配置,添加二元语法(bigram)特征,让模型能识别连续词对。此方法需手动调整配置文件,复杂度较高,适合不想引入Transformer的场景。


内容的提问来源于stack exchange,提问作者jacmkno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 05:31:09