如何从Stanza Constituency Parse Tree获取原字符串中的Token位置?
Stanza提取名词短语(NP)后映射原句位置的解决思路
问题背景
我使用Stanza提取文本中的名词短语(NP),并按其在句法树中的深度进行存储,代码如下:
from collections import defaultdict import stanza nlp = stanza.Pipeline('en', tokenize_pretokenized=True) sentence_tokens = ['This', 'is', 'a', 'sentence', '.'] doc = nlp(sentence_tokens) for sent in doc.sentences: tree = sent.constituency def extract_NPs(tree, np_dict): for child in tree.children: if child.label=='NP': np_dict[child.depth()].append(child) np_dict = extract_NPs(child, np_dict) return np_dict nps = extract_NPs(tree, np_dict=defaultdict(list))
输出为以深度为键、对应NP树列表为值的字典。当前遇到的核心问题是:无法将NP文本映射回原输入句子的对应位置,由于句子中存在大量重复Token,直接通过Token文本查找索引的方法完全不可行。
解决思路
1. 利用Stanza Tree内置的leaf_indices()方法
Stanza的Constituency Tree对象自带leaf_indices()方法,该方法直接返回当前子树所覆盖的所有叶子节点(即原句中的Token)在整个句子中的原始索引,完全规避了重复Token的匹配问题:
- 对于任意NP树,
tree.leaf_indices()会返回一个整数列表,包含该NP所有Token的原句索引 - 起始索引为列表第一个元素,结束索引为最后一个元素+1(适配Python左闭右开的切片规则)
2. 修改提取函数,同步记录索引范围
直接修改原有的extract_NPs函数,将每个NP对应的索引范围与NP树一起存储,后续使用时可直接定位原句位置。修改后的完整代码如下:
from collections import defaultdict import stanza nlp = stanza.Pipeline('en', tokenize_pretokenized=True) sentence_tokens = ['This', 'is', 'a', 'sentence', '.'] doc = nlp(sentence_tokens) for sent in doc.sentences: tree = sent.constituency def extract_NPs_with_indices(tree, np_dict): for child in tree.children: if child.label == 'NP': # 获取当前NP覆盖的Token原句索引 leaf_ids = child.leaf_indices() start_idx = leaf_ids[0] end_idx = leaf_ids[-1] + 1 # 存储NP树、起始索引、结束索引 np_dict[child.depth()].append( (child, start_idx, end_idx) ) np_dict = extract_NPs_with_indices(child, np_dict) return np_dict # 得到带索引的NP字典 nps_with_indices = extract_NPs_with_indices(tree, np_dict=defaultdict(list)) # 示例输出验证 for depth, items in nps_with_indices.items(): print(f"深度 {depth}:") for np_tree, start, end in items: np_text = ' '.join(np_tree.leaves()) original_tokens = sentence_tokens[start:end] print(f"NP文本: {np_text}, 原句索引范围: [{start}, {end}), 对应原Token: {original_tokens}]")
3. 重复Token场景的验证
若原句为['cat', 'chased', 'the', 'cat'],提取到的两个NP(第一个cat和第四个cat)会通过leaf_indices()得到不同的索引([0]和[3]),可精准区分重复Token的位置,完全解决原问题。
内容的提问来源于stack exchange,提问作者kachap
相关产品推荐
相关产品推荐

