如何在Colab中用Stanza获取CoreNLP enhanced++ dependencies
Google Colab环境下通过Stanza获取enhanced++依存关系操作步骤
你之前获取到的基础依存是Stanza原生神经解析器的默认输出,Stanza本身的Python封装已经内置了CoreNLP客户端能力,不需要单独部署CoreNLP服务,按以下步骤操作即可拿到enhanced++ dependencies:
1. 安装依赖
在Colab笔记本单元格中运行以下命令安装Stanza:
!pip install stanza import stanza
2. 下载CoreNLP运行包与对应语言模型
首次运行需要拉取CoreNLP运行环境和对应语言的解析模型,以英文为例:
# 拉取CoreNLP核心运行包 stanza.install_corenlp() # 下载英文解析模型,处理其他语言替换为对应语言标识即可 stanza.download_corenlp_models(model='english', version='4.5.4')
3. 配置并启动CoreNLP客户端
启动客户端时必须显式开启enhanced++依存输出配置,注意不要用stanza.Pipeline()初始化的原生神经pipeline,该pipeline目前仅支持输出基础依存:
from stanza.server import CoreNLPClient # 上下文管理器会自动启动、关闭CoreNLP服务 with CoreNLPClient( annotators=['tokenize', 'ssplit', 'pos', 'lemma', 'parse', 'depparse'], properties={ # 该参数指定输出完整enhanced++依存关系 'depparse.extradependencies': 'MAXIMAL', 'parse.model': 'edu/stanford/nlp/models/parser/nndep/english_UD.gz' }, memory='4G', endpoint='http://localhost:9000', be_quiet=True ) as client: # 替换为待解析的目标文本 target_text = "Replace this with the text you need to parse." annotation_result = client.annotate(target_text)
4. 提取enhanced++依存结果
解析结果的enhancedPlusPlusDependencies字段存储了完整的enhanced++依存边信息,可直接遍历提取:
for sentence in annotation_result['sentences']: token_list = [token['word'] for token in sentence['tokens']] print("当前句子Token序列:", token_list) print("Enhanced++ 依存关系:") for dep_edge in sentence['enhancedPlusPlusDependencies']: gov_pos = dep_edge['governor'] dep_pos = dep_edge['dependent'] rel_type = dep_edge['dep'] gov_token = token_list[gov_pos - 1] if gov_pos > 0 else 'ROOT' dep_token = token_list[dep_pos - 1] print(f"支配词:{gov_token}(位置{gov_pos}) → 依存词:{dep_token}(位置{dep_pos}) 关系类型:{rel_type}")
注意事项
- 首次运行会自动下载约1G的CoreNLP相关文件,需等待下载完成后再执行后续解析操作,Colab默认预装Java环境,无需额外配置Java运行环境
depparse.extradependencies参数不要填错:MAXIMAL对应完整enhanced++输出,COLLAPSED对应折叠后的基础依存输出- 如果需要处理中文等其他语言,替换对应语言模型即可,参数配置逻辑一致
内容的提问来源于stack exchange,提问作者Atharva Swami
相关产品推荐
相关产品推荐

