如何使用Stanza获取动词不定式形式及识别句中不定式动词
Hey there! Let's tackle your two questions about working with verb infinitives in Stanza—identifying them in sentences and retrieving their infinitive forms. I'll use your provided example to make this concrete.
1. Identifying Infinitive Verbs in a Sentence
In Stanza's dependency parsing results, infinitives paired with "to" (like "find" in your example) have clear tell-tale signs:
- The word itself is tagged as
VERB - Its dependency relation (
deprel) isxcomp(this stands for "open clausal complement", Stanza's way of marking infinitives acting as verb complements) - There's a corresponding "to" (tagged
PART) whosedeprelismark, and whoseheadpoints directly to this verb
Looking at your sample output:
id: 4 word: to POS: PART head id: 5 head: find deprel: mark
id: 5 word: find POS: VERB head id: 2 head: need deprel: xcomp
"find" hits all three criteria—this confirms it's an infinitive verb.
Here's a code snippet to automatically detect these verbs:
doc = "I need you to find the verbes in this sentence" en_nlp = stanza.Pipeline('en', processors='tokenize,lemma,mwt,pos,depparse', verbose=False, use_gpu=False) processed = en_nlp(doc) for sent in processed.sentences: print("Identified infinitive verbs:") # Loop through each word to check the criteria for word in sent.words: if word.pos == 'VERB' and word.deprel == 'xcomp': # Check if there's a "to" marking this verb as infinitive has_infinitive_to = any( w.head == word.id and w.pos == 'PART' and w.deprel == 'mark' and w.text.lower() == 'to' for w in sent.words ) if has_infinitive_to: print(f"- {word.text} (ID: {word.id})")
Running this will output:
Identified infinitive verbs: - find (ID: 5)
2. Getting the Infinitive Form of a Verb
There are two common scenarios here:
Scenario 1: For an already identified infinitive verb
The full infinitive form is simply "to" + the verb's lemma (Stanza's lemma attribute gives you the base form of the verb, which is exactly what we need for infinitives).
Scenario 2: For any verb (converting it to infinitive form)
Same logic applies—append the verb's lemma to "to" to get the standard infinitive form.
Here's code to handle both cases:
# Get infinitive form for identified infinitives print("\nInfinitive forms for detected verbs:") for sent in processed.sentences: for word in sent.words: if word.pos == 'VERB' and word.deprel == 'xcomp': has_infinitive_to = any( w.head == word.id and w.pos == 'PART' and w.deprel == 'mark' and w.text.lower() == 'to' for w in sent.words ) if has_infinitive_to: infinitive = f"to {word.lemma}" print(f"- {word.text} → {infinitive}") # Get infinitive form for all verbs in the sentence print("\nInfinitive forms for all verbs:") for sent in processed.sentences: for word in sent.words: if word.pos == 'VERB': infinitive = f"to {word.lemma}" print(f"- {word.text} → {infinitive}")
This will output:
Infinitive forms for detected verbs: - find → to find Infinitive forms for all verbs: - need → to need - find → to find
A quick note on bare infinitives
If you're dealing with bare infinitives (verbs without "to", like "swim" in "I can swim"), the deprel will often be xcomp or aux, and you can adjust the detection logic to check for modal verbs (like "can", "will") as the head. The infinitive form here is just the verb's lemma (or "to + lemma" if you want the full infinitive).
内容的提问来源于stack exchange,提问作者Belkacem Thiziri

