Pandas存储多级索引数据为float32型Parquet时索引精度丢失如何解决
问题根因
你当前的脚本在遍历table.schema批量转换字段类型时,没有区分普通数据列和存储索引值的元数据字段,将索引的number层级对应的float64字段也统一转换成了float32,才会出现精度损失。
推荐解决方案(最简单不易出错)
直接在Pandas层完成数据列的类型转换,索引和多级列的类型不会被修改,不需要手动操作PyArrow schema:
import pandas as pd import pyarrow as pa import pyarrow.parquet as pq df = pd.DataFrame({"col1": [1.0, 2.0, 3.0], "col2": [2.3, 2.4, 2.5], "col3": [3.1, 3.2, 3.3]}) df.index = pd.MultiIndex.from_tuples([('a', '2021-01-01', 100), ('a', '2021-01-01', 200), ('a', '2021-01-01', 7080.39)], names=('name', 'date', 'number')) df.columns = pd.MultiIndex.from_tuples([('a', '2021-01-01', 100), ('a', '2021-01-01', 200), ('a', '2021-01-01', 7080.39)], names=('name', 'date', 'number')) # 仅转换数据列为float32,索引、列的类型完全不变 df = df.astype('float32') # 直接写入Parquet即可 table = pa.Table.from_pandas(df) pq.write_table(table, 'float.parquet') # 读取验证 df2 = pd.read_parquet('float.parquet') print(df2.index) # 输出的number层级为准确的7080.39,无精度损失
备选方案(PyArrow Schema层处理)
如果你坚持要在PyArrow层面修改schema,可以先解析元数据拿到索引字段列表,转换时跳过这些字段:
import json table = pa.Table.from_pandas(df) # 从元数据中提取索引字段名 pandas_meta = json.loads(table.schema.metadata[b'pandas'].decode()) index_col_names = {col['name'] for col in pandas_meta['index_columns']} # 构造schema时仅转换非索引的float64字段 schema = pa.schema([ pa.field(field.name, pa.float32() if (field.type == pa.float64() and field.name not in index_col_names) else field.type) for field in table.schema ], metadata=table.schema.metadata) table = table.cast(schema) pq.write_table(table, 'float.parquet')
内容的提问来源于stack exchange,提问作者abisko
相关产品推荐
相关产品推荐

