将scispacy Hearst Patterns返回的Span对象存入DataFrame时出错
scispacy Hearst Patterns结果存入DataFrame触发长度不匹配错误
使用scispacy的Hearst Patterns功能返回的spacy.tokens.span.Span对象,尝试存入pandas DataFrame时触发ValueError: All arrays must be of the same length错误,核心问题是Span对象被拆分为多个词,导致各字段长度不一致。
示例代码
import spacy from scispacy.hyponym_detector import HyponymDetector nlp = spacy.load("en_core_sci_sm") nlp.add_pipe("hyponym_detector", last=True, config={"extended": False}) doc = nlp("Keystone plant species such as fig trees are good for the soil.") print(doc._.hearst_patterns) >>> [('such_as', Keystone plant species, fig trees)] print(type(doc_hp[0][1])) >>> <class 'spacy.tokens.span.Span'> doc_hp = doc._.hearst_patterns dict = { "hp_connector": doc_hp[0][0], "hp_entity_1":doc_hp[0][1], "hp_entity_2":doc_hp[0][2], } df = pd.DataFrame.from_dict(dict)
错误信息
Traceback (most recent call last): File "extract_hearst_patterns.py", line 42, in <module> df = pd.DataFrame.from_dict(dict) File "/venv/lib/python3.9/site-packages/pandas/core/frame.py", line 1760, in from_dict return cls(data, index=index, columns=columns, dtype=dtype) File "/venv/lib/python3.9/site-packages/pandas/core/frame.py", line 709, in __init__ mgr = dict_to_mgr(data, index, columns, dtype=dtype, copy=copy, typ=manager) File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 481, in dict_to_mgr return arrays_to_mgr(arrays, columns, index, dtype=dtype, typ=typ, consolidate=copy) File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 115, in arrays_to_mgr index = _extract_index(arrays) File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 655, in _extract_index raise ValueError("All arrays must be of the same length") ValueError: All arrays must be of the same length
错误原因
pandas会将spacy.tokens.span.Span对象视为可迭代的token序列(Span本质是多个token的集合),因此hp_entity_1和hp_entity_2会被拆分成单个token的列表,而hp_connector是单个字符串,三者长度不匹配,触发错误。
解决方法
方法1:将Span转为字符串(最常用)
把Span对象直接转为文本字符串,确保每个字段都是单个值,长度一致:
import spacy import pandas as pd from scispacy.hyponym_detector import HyponymDetector nlp = spacy.load("en_core_sci_sm") nlp.add_pipe("hyponym_detector", last=True, config={"extended": False}) doc = nlp("Keystone plant species such as fig trees are good for the soil.") doc_hp = doc._.hearst_patterns data_dict = { "hp_connector": [doc_hp[0][0]], "hp_entity_1": [str(doc_hp[0][1])], "hp_entity_2": [str(doc_hp[0][2])], } df = pd.DataFrame.from_dict(data_dict) print(df)
输出:
hp_connector hp_entity_1 hp_entity_2 0 such_as Keystone plant species fig trees
方法2:保留Span对象
如果需要后续使用Span的NLP属性(比如token的词性、实体类型),要把每个字段放在列表中,让pandas识别为单行数据:
data_dict = { "hp_connector": [doc_hp[0][0]], "hp_entity_1": [doc_hp[0][1]], "hp_entity_2": [doc_hp[0][2]], } df = pd.DataFrame.from_dict(data_dict) # 验证类型 print(type(df.loc[0, "hp_entity_1"])) # <class 'spacy.tokens.span.Span'>
方法3:批量处理多个结果
如果文本中有多个Hearst Pattern结果,用列表推导式批量转换:
data = [] for pattern in doc_hp: data.append({ "hp_connector": pattern[0], "hp_entity_1": str(pattern[1]), "hp_entity_2": str(pattern[2]) }) df = pd.DataFrame(data)
内容的提问来源于stack exchange,提问作者learningtocode
相关产品推荐
相关产品推荐

