You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法获取Berkeley Neural Parser终端标签的技术咨询

使用Benepar获取成分标签时缺失终端标签的问题

问题代码

import spacy, benepar 

nlp = spacy.load("en_core_web_sm")
nlp.add_pipe('benepar', config={'model': 'benepar_en3'})
doc = nlp('This red apple is tasty')

sent = list(doc.sents)[0]
tree = sent._.parse_string

for child in sent._.constituents:
    print(child)
    print(child._.labels)

问题描述

运行上述代码后,多个节点的._.labels属性为空,缺少DT、JJ这类终端标签,但Benepar测试网站的解析结果包含这些标签,询问是否是._.labels属性使用有误,或是否有其他获取途径。

运行时警告信息(翻译后)

  • 你正在使用<class 'transformers.models.t5.tokenization_t5.T5Tokenizer'>的旧版行为,这意味着特殊token之后的token无法被正确处理。建议阅读相关PR。
  • 你正在使用T5TokenizerFast分词器,请注意使用__call__方法比先编码文本再调用pad方法获取填充编码更快。
  • /home/mllenouvelle/.local/lib/python3.10/site-packages/torch/distributions/distribution.py:51: UserWarning: <class 'torch_struct.distributions.TreeCRF'>未定义arg_constraints,请设置arg_constraints = {}或使用validate_args=False初始化分布以关闭验证。

解决方法

sent._.constituents仅返回短语级的非终端节点,不会包含单个token对应的终端节点(比如DT、JJ标签对应的单词),因此看不到这类标签。要获取所有节点(包括终端节点),需要递归遍历整个解析树:

修正后的代码

import spacy, benepar

nlp = spacy.load("en_core_web_sm")
nlp.add_pipe('benepar', config={'model': 'benepar_en3'})
doc = nlp('This red apple is tasty')
sent = list(doc.sents)[0]

# 递归遍历解析树的所有节点
def traverse_parse_tree(node):
    print(f"文本: {node.text}, 标签: {node._.labels}")
    # 遍历所有子节点(包括终端和非终端)
    if node._.children:
        for child in node._.children:
            traverse_parse_tree(child)

traverse_parse_tree(sent)

说明

  • node._.children会返回当前节点的所有子节点,既包含短语级的非终端节点,也包含单个token的终端节点。
  • 单个token节点的._.labels会返回对应的POS标签(如DT、JJ),短语节点则返回成分标签(如NP、VP)。

关于警告信息的补充:

  • 前两条是transformers分词器的版本提示,不影响标签获取功能,若要消除可升级transformers至支持新版T5Tokenizer的版本。
  • 第三条是torch_struct的警告,同样不影响核心功能,可直接忽略,或按提示设置参数关闭警告。

内容的提问来源于stack exchange,提问作者Evan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 08:13:11