You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将scispacy Hearst Patterns返回的Span对象存入DataFrame时出错

scispacy Hearst Patterns结果存入DataFrame触发长度不匹配错误

使用scispacy的Hearst Patterns功能返回的spacy.tokens.span.Span对象,尝试存入pandas DataFrame时触发ValueError: All arrays must be of the same length错误,核心问题是Span对象被拆分为多个词,导致各字段长度不一致。

示例代码

import spacy
from scispacy.hyponym_detector import HyponymDetector

nlp = spacy.load("en_core_sci_sm")
nlp.add_pipe("hyponym_detector", last=True, config={"extended": False})

doc = nlp("Keystone plant species such as fig trees are good for the soil.")

print(doc._.hearst_patterns)
>>> [('such_as', Keystone plant species, fig trees)]
print(type(doc_hp[0][1]))
>>> <class 'spacy.tokens.span.Span'>

doc_hp = doc._.hearst_patterns
dict = {
    "hp_connector": doc_hp[0][0],
    "hp_entity_1":doc_hp[0][1],
    "hp_entity_2":doc_hp[0][2],
}

df = pd.DataFrame.from_dict(dict)

错误信息

Traceback (most recent call last):
  File "extract_hearst_patterns.py", line 42, in <module>
    df = pd.DataFrame.from_dict(dict)
  File "/venv/lib/python3.9/site-packages/pandas/core/frame.py", line 1760, in from_dict
    return cls(data, index=index, columns=columns, dtype=dtype)
  File "/venv/lib/python3.9/site-packages/pandas/core/frame.py", line 709, in __init__
    mgr = dict_to_mgr(data, index, columns, dtype=dtype, copy=copy, typ=manager)
  File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 481, in dict_to_mgr
    return arrays_to_mgr(arrays, columns, index, dtype=dtype, typ=typ, consolidate=copy)
  File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 115, in arrays_to_mgr
    index = _extract_index(arrays)
  File "/venv/lib/python3.9/site-packages/pandas/core/internals/construction.py", line 655, in _extract_index
    raise ValueError("All arrays must be of the same length")
ValueError: All arrays must be of the same length

错误原因

pandas会将spacy.tokens.span.Span对象视为可迭代的token序列(Span本质是多个token的集合),因此hp_entity_1和hp_entity_2会被拆分成单个token的列表,而hp_connector是单个字符串,三者长度不匹配,触发错误。

解决方法

方法1:将Span转为字符串(最常用)

把Span对象直接转为文本字符串,确保每个字段都是单个值,长度一致:

import spacy
import pandas as pd
from scispacy.hyponym_detector import HyponymDetector

nlp = spacy.load("en_core_sci_sm")
nlp.add_pipe("hyponym_detector", last=True, config={"extended": False})

doc = nlp("Keystone plant species such as fig trees are good for the soil.")
doc_hp = doc._.hearst_patterns

data_dict = {
    "hp_connector": [doc_hp[0][0]],
    "hp_entity_1": [str(doc_hp[0][1])],
    "hp_entity_2": [str(doc_hp[0][2])],
}

df = pd.DataFrame.from_dict(data_dict)
print(df)

输出:

hp_connector          hp_entity_1 hp_entity_2
0      such_as  Keystone plant species  fig trees

方法2:保留Span对象

如果需要后续使用Span的NLP属性(比如token的词性、实体类型),要把每个字段放在列表中,让pandas识别为单行数据:

data_dict = {
    "hp_connector": [doc_hp[0][0]],
    "hp_entity_1": [doc_hp[0][1]],
    "hp_entity_2": [doc_hp[0][2]],
}

df = pd.DataFrame.from_dict(data_dict)
# 验证类型
print(type(df.loc[0, "hp_entity_1"]))  # <class 'spacy.tokens.span.Span'>

方法3:批量处理多个结果

如果文本中有多个Hearst Pattern结果,用列表推导式批量转换:

data = []
for pattern in doc_hp:
    data.append({
        "hp_connector": pattern[0],
        "hp_entity_1": str(pattern[1]),
        "hp_entity_2": str(pattern[2])
    })

df = pd.DataFrame(data)

内容的提问来源于stack exchange,提问作者learningtocode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 06:43:35