You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Stanza Constituency Parse Tree获取原字符串中的Token位置?

Stanza提取名词短语(NP)后映射原句位置的解决思路

问题背景

我使用Stanza提取文本中的名词短语(NP),并按其在句法树中的深度进行存储,代码如下:

from collections import defaultdict
import stanza

nlp = stanza.Pipeline('en', tokenize_pretokenized=True)
sentence_tokens = ['This', 'is', 'a', 'sentence', '.']
doc = nlp(sentence_tokens)
for sent in doc.sentences:
    tree = sent.constituency

    def extract_NPs(tree, np_dict):
        for child in tree.children:
            if child.label=='NP':
                np_dict[child.depth()].append(child)
            np_dict = extract_NPs(child, np_dict)
        return np_dict
    nps = extract_NPs(tree, np_dict=defaultdict(list))

输出为以深度为键、对应NP树列表为值的字典。当前遇到的核心问题是:无法将NP文本映射回原输入句子的对应位置,由于句子中存在大量重复Token,直接通过Token文本查找索引的方法完全不可行。


解决思路

1. 利用Stanza Tree内置的leaf_indices()方法

Stanza的Constituency Tree对象自带leaf_indices()方法,该方法直接返回当前子树所覆盖的所有叶子节点(即原句中的Token)在整个句子中的原始索引,完全规避了重复Token的匹配问题:

  • 对于任意NP树,tree.leaf_indices()会返回一个整数列表,包含该NP所有Token的原句索引
  • 起始索引为列表第一个元素,结束索引为最后一个元素+1(适配Python左闭右开的切片规则)

2. 修改提取函数,同步记录索引范围

直接修改原有的extract_NPs函数,将每个NP对应的索引范围与NP树一起存储,后续使用时可直接定位原句位置。修改后的完整代码如下:

from collections import defaultdict
import stanza

nlp = stanza.Pipeline('en', tokenize_pretokenized=True)
sentence_tokens = ['This', 'is', 'a', 'sentence', '.']
doc = nlp(sentence_tokens)

for sent in doc.sentences:
    tree = sent.constituency

    def extract_NPs_with_indices(tree, np_dict):
        for child in tree.children:
            if child.label == 'NP':
                # 获取当前NP覆盖的Token原句索引
                leaf_ids = child.leaf_indices()
                start_idx = leaf_ids[0]
                end_idx = leaf_ids[-1] + 1
                # 存储NP树、起始索引、结束索引
                np_dict[child.depth()].append( (child, start_idx, end_idx) )
            np_dict = extract_NPs_with_indices(child, np_dict)
        return np_dict
    
    # 得到带索引的NP字典
    nps_with_indices = extract_NPs_with_indices(tree, np_dict=defaultdict(list))

# 示例输出验证
for depth, items in nps_with_indices.items():
    print(f"深度 {depth}:")
    for np_tree, start, end in items:
        np_text = ' '.join(np_tree.leaves())
        original_tokens = sentence_tokens[start:end]
        print(f"NP文本: {np_text}, 原句索引范围: [{start}, {end}), 对应原Token: {original_tokens}]")

3. 重复Token场景的验证

若原句为['cat', 'chased', 'the', 'cat'],提取到的两个NP(第一个cat和第四个cat)会通过leaf_indices()得到不同的索引([0]和[3]),可精准区分重复Token的位置,完全解决原问题。


内容的提问来源于stack exchange,提问作者kachap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 00:30:38