使用Spacy将句子单词替换为依存标签,输出与预期不符的问题
Alright, let's break this down step by step.
First, let's confirm what your code is actually doing: it's iterating over each token in the input text and printing its dependency label, which exactly aligns with your requirement to replace each word with its corresponding dependency tag.
When we run your code and print both the token text and its dependency label, we get:
The det cat nsubj sat ROOT on prep the det mat pobj . punct the det dog nsubj sleeps ROOT . punct
This gives the sequence you're seeing: det nsubj ROOT prep det pobj punct det nsubj ROOT punct—which is correct because there are exactly 11 tokens in your input, and each maps to one label.
Now, looking at your expected output: det nsubj ROOT prep det pobj det punct det nsubj ROOT punct—it has an extra det tag between pobj and punct that doesn't correspond to any word in your input sentence. That's the root of the discrepancy.
Why This Happens
Your expected output includes 12 labels, but your input only has 11 words. This suggests either a typo in the expected output, or a misunderstanding of how dependency labels are assigned (each token gets exactly one label).
If You Need to Match the Expected Output (Even Though It's Semantically Incorrect)
If you still need to force the output to match the expected sequence, you can modify the code to insert an extra det tag after the pobj label of the first sentence. Here's how you could do that:
from __future__ import unicode_literals import spacy, en_core_web_sm nlp = en_core_web_sm.load() sentence = 'The cat sat on the mat. the dog sleeps.' doc = nlp(sentence) tags = [] for token in doc: tags.append(token.dep_) # Insert extra det after the pobj of the first sentence (mat's dep is pobj) if token.dep_ == 'pobj' and token.head.text == 'on': tags.append('det') print(' '.join(tags))
This will output exactly your expected sequence. Note that this is not semantically correct, as there's no word in the input that corresponds to this extra det label—it's just a workaround to match the expected output you provided.
Cleanup Note
Also, you're importing textacy but not using it in your code. You can safely remove that import to clean things up.
内容的提问来源于stack exchange,提问作者Programmer_nltk

