You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Spacy将句子单词替换为依存标签,输出与预期不符的问题

Fixing the Dependency Tag Sequence Discrepancy

Alright, let's break this down step by step.

First, let's confirm what your code is actually doing: it's iterating over each token in the input text and printing its dependency label, which exactly aligns with your requirement to replace each word with its corresponding dependency tag.

When we run your code and print both the token text and its dependency label, we get:

The det
cat nsubj
sat ROOT
on prep
the det
mat pobj
. punct
the det
dog nsubj
sleeps ROOT
. punct

This gives the sequence you're seeing: det nsubj ROOT prep det pobj punct det nsubj ROOT punct—which is correct because there are exactly 11 tokens in your input, and each maps to one label.

Now, looking at your expected output: det nsubj ROOT prep det pobj det punct det nsubj ROOT punct—it has an extra det tag between pobj and punct that doesn't correspond to any word in your input sentence. That's the root of the discrepancy.

Why This Happens

Your expected output includes 12 labels, but your input only has 11 words. This suggests either a typo in the expected output, or a misunderstanding of how dependency labels are assigned (each token gets exactly one label).

If You Need to Match the Expected Output (Even Though It's Semantically Incorrect)

If you still need to force the output to match the expected sequence, you can modify the code to insert an extra det tag after the pobj label of the first sentence. Here's how you could do that:

from __future__ import unicode_literals
import spacy, en_core_web_sm

nlp = en_core_web_sm.load()
sentence = 'The cat sat on the mat. the dog sleeps.'
doc = nlp(sentence)

tags = []
for token in doc:
    tags.append(token.dep_)
    # Insert extra det after the pobj of the first sentence (mat's dep is pobj)
    if token.dep_ == 'pobj' and token.head.text == 'on':
        tags.append('det')

print(' '.join(tags))

This will output exactly your expected sequence. Note that this is not semantically correct, as there's no word in the input that corresponds to this extra det label—it's just a workaround to match the expected output you provided.

Cleanup Note

Also, you're importing textacy but not using it in your code. You can safely remove that import to clean things up.

内容的提问来源于stack exchange,提问作者Programmer_nltk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:31:41