You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame单元格中的列表保存为HDF5表格格式?

问题:含列表列的DataFrame无法以表格格式保存到HDF5

场景与复现

DataFrame结构:

column1
0   [0, 1, 2, 3, 4]

复现代码:

import pandas as pd
test = pd.DataFrame({"column1":[list(range(0,5))]})
test.to_hdf('test','testgroup',format="table")

执行后报错:

---------------------------------------------------------------------------

TypeError                                 Traceback (most recent call last)

<ipython-input-65-c2dbeaca15df> in <module>
      1 test = pd.DataFrame({"column1":[list(range(0,5))]})
----> 2 test.to_hdf('test','testgroup',format="table")

7 frames

/usr/local/lib/python3.7/dist-packages/pandas/io/pytables.py in _maybe_convert_for_string_atom(name, block, existing_col, min_itemsize, nan_rep, encoding, errors, columns)
   4979                 error_column_label = columns[i] if len(columns) > i else f"No.{i}"
   4980                 raise TypeError(
-> 4981                     f"Cannot serialize the column [{error_column_label}]\n"
   4982                     f"because its data contents are not [string] but "
   4983                     f"[{inferred_type}] object dtype"

TypeError: Cannot serialize the column [column1]
because its data contents are not [string] but [mixed] object dtype

已知拆分列表到多列、转字符串再还原的方案不适用,需找到直接将列表保存到HDF5表格格式且支持追加的标准方法。


解决方案

方法1:Pickle序列化列后保存

HDF5表格格式仅支持字符串类型的object列,可先将列表用pickle序列化为字节,保存后再反序列化还原:

import pandas as pd
import pickle

# 序列化列表列
test = pd.DataFrame({"column1":[list(range(0,5))]})
test['column1'] = test['column1'].apply(pickle.dumps)

# 以表格格式追加保存
test.to_hdf('test.h5', 'testgroup', format='table', append=True)

# 读取并反序列化
df_read = pd.read_hdf('test.h5', 'testgroup')
df_read['column1'] = df_read['column1'].apply(pickle.loads)
print(df_read)

方法2:直接使用PyTables操作可变长度数组

借助PyTables底层API,创建支持可变长度数组的表格,原生适配列表存储:

import pandas as pd
import tables as tb

# 定义可变长度整数数组类型
var_int_array = tb.VLArrayAtom(dtype=tb.Int64Atom())

# 创建HDF5文件与表格
with tb.open_file('test_vl.h5', 'w') as f:
    table = f.create_table(f.root, 'testgroup', 
                          {'column1': var_int_array},
                          title='Table with variable-length arrays')
    
    # 添加数据(支持循环批量追加)
    row = table.row
    row['column1'] = list(range(5))
    row.append()
    table.flush()

# 读取数据并转为DataFrame
with tb.open_file('test_vl.h5', 'r') as f:
    table = f.root.testgroup
    data = [row['column1'] for row in table]
    df_read = pd.DataFrame({'column1': data})
print(df_read)

方法3:固定格式保存(不支持追加,仅作参考)

若无需追加功能,format="fixed"可直接保存含列表的DataFrame,但无法后续追加数据:

test.to_hdf('test_fixed.h5', 'testgroup', format='fixed')
df_read = pd.read_hdf('test_fixed.h5', 'testgroup')

内容的提问来源于stack exchange,提问作者Andrei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 02:20:55