You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spacy中Span转Doc丢失Benepar句法分析数据的解决方法

问题

我有一个包含多个句子的字符串,想要获取每个句子的constituency parse。先通过nlp解析整个字符串得到Spacy Doc对象,再遍历doc.sents并通过span.as_doc()将Span转换为Doc,但转换后原有的benepar constituency parse数据丢失了。测试代码如下:

import spacy
import benepar

nlp = spacy.load("en_core_sci_md", disable=["ner", "lemmatizer", "textcat"])
nlp.add_pipe('benepar', config={'model': BENEPAR_DIR})
nlp_test1 = nlp('The quick brown fox jumps over the lazy dog')
print(list(nlp_test1.sents)[0]._.parse_string) # 正常输出benepar解析结果

nlp_test2 = list(nlp_test1.sents)[0].as_doc()
print(list(nlp_test2.sents)[0]._.parse_string) # 找不到constituency parse数据

# 尝试传递array_attrs也无效
nlp_test3 = list(nlp_test1.sents)[0].as_doc(array_head=nlp_test1._get_array_attrs())
print(list(nlp_test3.sents)[0]._.parse_string)

请问如何在将Span转换为Doc时保留benepar constituency parse数据?或者benepar仅能解析doc.sents中的第一个句子?

解决方案

benepar可以解析所有句子,问题出在as_doc()转换时,Spacy不会自动将自定义扩展属性(比如benepar的_.parse_string)复制到新生成的Doc对象中。以下是两种可行的解决方法:

方法一:直接访问原Span的解析结果(推荐)

不需要将Span转换为新Doc,直接遍历原Doc的doc.sents,对每个Sentence Span直接调用_.parse_string即可获取解析结果:

nlp_test1 = nlp('Sentence one. Sentence two.')
for sent in nlp_test1.sents:
    print(sent._.parse_string)

这种方式无需额外转换,直接利用原解析结果,效率最高。

方法二:手动复制扩展属性或重新解析

如果确实需要将Span转为Doc并保留benepar数据,可选择以下两种方式:

1. 手动复制扩展属性

创建新Doc后,将原Span的_.parse_string赋值给新Doc对应范围的扩展属性:

sent_span = list(nlp_test1.sents)[0]
new_doc = sent_span.as_doc()
# 复制benepar解析结果到新Doc
new_doc[0:len(new_doc)]._.parse_string = sent_span._.parse_string
print(new_doc._.parse_string)

2. 重新运行benepar管道

对转换后的新Doc重新调用nlp管道,让benepar再次解析该文本:

sent_span = list(nlp_test1.sents)[0]
new_doc = sent_span.as_doc()
# 重新运行nlp管道(包含benepar)
nlp(new_doc)
print(new_doc._.parse_string)

注意:这种方式会重复解析文本,效率低于直接访问原Span的结果,仅在必要时使用。

另外,测试代码中的nlp_test是笔误,应改为nlp_test1。

内容的提问来源于stack exchange,提问作者Dan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 14:05:39