如何将Pandas DataFrame单元格中的列表保存为HDF5表格格式?
问题:含列表列的DataFrame无法以表格格式保存到HDF5
场景与复现
DataFrame结构:
column1 0 [0, 1, 2, 3, 4]
复现代码:
import pandas as pd test = pd.DataFrame({"column1":[list(range(0,5))]}) test.to_hdf('test','testgroup',format="table")
执行后报错:
--------------------------------------------------------------------------- TypeError Traceback (most recent call last) <ipython-input-65-c2dbeaca15df> in <module> 1 test = pd.DataFrame({"column1":[list(range(0,5))]}) ----> 2 test.to_hdf('test','testgroup',format="table") 7 frames /usr/local/lib/python3.7/dist-packages/pandas/io/pytables.py in _maybe_convert_for_string_atom(name, block, existing_col, min_itemsize, nan_rep, encoding, errors, columns) 4979 error_column_label = columns[i] if len(columns) > i else f"No.{i}" 4980 raise TypeError( -> 4981 f"Cannot serialize the column [{error_column_label}]\n" 4982 f"because its data contents are not [string] but " 4983 f"[{inferred_type}] object dtype" TypeError: Cannot serialize the column [column1] because its data contents are not [string] but [mixed] object dtype
已知拆分列表到多列、转字符串再还原的方案不适用,需找到直接将列表保存到HDF5表格格式且支持追加的标准方法。
解决方案
方法1:Pickle序列化列后保存
HDF5表格格式仅支持字符串类型的object列,可先将列表用pickle序列化为字节,保存后再反序列化还原:
import pandas as pd import pickle # 序列化列表列 test = pd.DataFrame({"column1":[list(range(0,5))]}) test['column1'] = test['column1'].apply(pickle.dumps) # 以表格格式追加保存 test.to_hdf('test.h5', 'testgroup', format='table', append=True) # 读取并反序列化 df_read = pd.read_hdf('test.h5', 'testgroup') df_read['column1'] = df_read['column1'].apply(pickle.loads) print(df_read)
方法2:直接使用PyTables操作可变长度数组
借助PyTables底层API,创建支持可变长度数组的表格,原生适配列表存储:
import pandas as pd import tables as tb # 定义可变长度整数数组类型 var_int_array = tb.VLArrayAtom(dtype=tb.Int64Atom()) # 创建HDF5文件与表格 with tb.open_file('test_vl.h5', 'w') as f: table = f.create_table(f.root, 'testgroup', {'column1': var_int_array}, title='Table with variable-length arrays') # 添加数据(支持循环批量追加) row = table.row row['column1'] = list(range(5)) row.append() table.flush() # 读取数据并转为DataFrame with tb.open_file('test_vl.h5', 'r') as f: table = f.root.testgroup data = [row['column1'] for row in table] df_read = pd.DataFrame({'column1': data}) print(df_read)
方法3:固定格式保存(不支持追加,仅作参考)
若无需追加功能,format="fixed"可直接保存含列表的DataFrame,但无法后续追加数据:
test.to_hdf('test_fixed.h5', 'testgroup', format='fixed') df_read = pd.read_hdf('test_fixed.h5', 'testgroup')
内容的提问来源于stack exchange,提问作者Andrei
相关产品推荐
相关产品推荐

