如何实现Polars数据类型的序列化与反序列化?
保存/读取Polars数据类型到文本文件的最优方案
直接序列化Polars Schema会报错,因为DataType对象默认不支持JSON序列化。以下是几种可靠的实现方案:
一、官方内置方案(推荐)
Polars的Schema类提供了原生的to_json()和from_json()方法,完美支持复杂数据类型(如List、Struct等)的序列化与反序列化。
保存Schema到JSON文件
import polars as pl # 示例DataFrame df = pl.DataFrame(schema={'a': pl.Int64, 'b': pl.List(pl.Float64)}) # 将Schema序列化为JSON字符串 schema_json = df.schema.to_json() # 写入文件 with open('schema.json', 'w') as f: f.write(schema_json)
从JSON文件恢复Schema
import polars as pl # 读取JSON文件内容 with open('schema.json', 'r') as f: schema_json = f.read() # 反序列化为Polars Schema restored_schema = pl.Schema.from_json(schema_json) # 使用恢复的Schema创建DataFrame或读取数据 new_df = pl.DataFrame(schema=restored_schema)
二、手动字符串转换方案(适合简单场景)
如果需要自定义格式,可以将数据类型转为字符串保存,读取时再解析回Polars类型。
保存Schema为字符串字典
import polars as pl import json df = pl.DataFrame(schema={'a': pl.Int64, 'b': pl.List(pl.Float64)}) # 将Schema转为字符串字典 schema_str_dict = {col: str(dtype) for col, dtype in df.schema.items()} # 写入JSON文件 with open('schema_str.json', 'w') as f: json.dump(schema_str_dict, f)
读取并解析字符串字典为Schema
import polars as pl import json # 读取JSON文件 with open('schema_str.json', 'r') as f: schema_str_dict = json.load(f) # 解析字符串为Polars数据类型 restored_schema = {col: pl.parse_dtype(dtype_str) for col, dtype_str in schema_str_dict.items()} # 创建DataFrame new_df = pl.DataFrame(schema=restored_schema)
三、CSV文件的Schema处理
CSV本身不存储Schema信息,因此需要单独保存Schema,读取CSV时手动指定:
保存数据与Schema
import polars as pl df = pl.DataFrame(schema={'a': pl.Int64, 'b': pl.List(pl.Float64)}, data=[(1, [1.0, 2.0])]) # 保存数据到CSV df.write_csv('data.csv') # 保存Schema到JSON文件 with open('schema.json', 'w') as f: f.write(df.schema.to_json())
读取CSV并应用Schema
import polars as pl # 恢复Schema with open('schema.json', 'r') as f: restored_schema = pl.Schema.from_json(f.read()) # 读取CSV时指定Schema,避免自动推断错误 df = pl.read_csv('data.csv', schema=restored_schema)
注意:优先使用官方内置的
to_json()/from_json()方法,它能可靠处理所有Polars数据类型,手动字符串转换仅适合简单类型场景。
内容的提问来源于stack exchange,提问作者Gabriel
相关产品推荐
相关产品推荐

