Spacy中Span转Doc丢失Benepar句法分析数据的解决方法
问题
我有一个包含多个句子的字符串,想要获取每个句子的constituency parse。先通过nlp解析整个字符串得到Spacy Doc对象,再遍历doc.sents并通过span.as_doc()将Span转换为Doc,但转换后原有的benepar constituency parse数据丢失了。测试代码如下:
import spacy import benepar nlp = spacy.load("en_core_sci_md", disable=["ner", "lemmatizer", "textcat"]) nlp.add_pipe('benepar', config={'model': BENEPAR_DIR}) nlp_test1 = nlp('The quick brown fox jumps over the lazy dog') print(list(nlp_test1.sents)[0]._.parse_string) # 正常输出benepar解析结果 nlp_test2 = list(nlp_test1.sents)[0].as_doc() print(list(nlp_test2.sents)[0]._.parse_string) # 找不到constituency parse数据 # 尝试传递array_attrs也无效 nlp_test3 = list(nlp_test1.sents)[0].as_doc(array_head=nlp_test1._get_array_attrs()) print(list(nlp_test3.sents)[0]._.parse_string)
请问如何在将Span转换为Doc时保留benepar constituency parse数据?或者benepar仅能解析doc.sents中的第一个句子?
解决方案
benepar可以解析所有句子,问题出在as_doc()转换时,Spacy不会自动将自定义扩展属性(比如benepar的_.parse_string)复制到新生成的Doc对象中。以下是两种可行的解决方法:
方法一:直接访问原Span的解析结果(推荐)
不需要将Span转换为新Doc,直接遍历原Doc的doc.sents,对每个Sentence Span直接调用_.parse_string即可获取解析结果:
nlp_test1 = nlp('Sentence one. Sentence two.') for sent in nlp_test1.sents: print(sent._.parse_string)
这种方式无需额外转换,直接利用原解析结果,效率最高。
方法二:手动复制扩展属性或重新解析
如果确实需要将Span转为Doc并保留benepar数据,可选择以下两种方式:
1. 手动复制扩展属性
创建新Doc后,将原Span的_.parse_string赋值给新Doc对应范围的扩展属性:
sent_span = list(nlp_test1.sents)[0] new_doc = sent_span.as_doc() # 复制benepar解析结果到新Doc new_doc[0:len(new_doc)]._.parse_string = sent_span._.parse_string print(new_doc._.parse_string)
2. 重新运行benepar管道
对转换后的新Doc重新调用nlp管道,让benepar再次解析该文本:
sent_span = list(nlp_test1.sents)[0] new_doc = sent_span.as_doc() # 重新运行nlp管道(包含benepar) nlp(new_doc) print(new_doc._.parse_string)
注意:这种方式会重复解析文本,效率低于直接访问原Span的结果,仅在必要时使用。
另外,测试代码中的nlp_test是笔误,应改为nlp_test1。
内容的提问来源于stack exchange,提问作者Dan
相关产品推荐
相关产品推荐

