使用Polars创建DataFrame遇ComputeError报错,Pandas可正常运行求解
Polars创建DataFrame时触发ComputeError的问题
错误信息
ComputeError: could not append value: 1.41431 of type: f64 to the builder; make sure that all rows have the same schema or consider increasing infer_schema_length it might also be that a value overflows the data-type's capacity
复现代码
import json import pandas as pd import polars as pl import requests # 加载测试JSON数据 response = requests.get('测试JSON文件地址') res = json.loads(response.text) # Pandas创建DataFrame正常运行 df_pd = pd.DataFrame(res) # Polars创建DataFrame抛出上述错误 df = pl.DataFrame(res)
问题说明
原本预期pl.DataFrame(res)能和pd.DataFrame(res)一样正常执行,但Polars抛出了上述错误。使用的是最新版Polars(0.20.31),曾参考过Polars的一个已关闭议题,但问题仍未解决。
已尝试的操作均无效:
- 检查过res中值的类型,表面看起来一致
- 将浮点数保留5位小数
- 增大
infer_schema_length参数
解决思路
1. 强制指定Schema
Polars的类型推断比Pandas严格,哪怕某一列前N行是整数,后面出现浮点数就会触发错误。直接手动指定每列的数据类型,跳过自动推断:
# 根据实际数据结构定义schema,示例如下 schema = { "column1": pl.Float64, "column2": pl.String, # 其他列依次对应指定类型 } df = pl.DataFrame(res, schema=schema)
2. 预处理统一数据类型
遍历所有行,把数值类型统一转为float64,避免混合类型导致推断失败:
processed_data = [] for row in res: processed_row = {} for key, value in row.items(): # 把整数、浮点数统一转为float64 if isinstance(value, (int, float)): processed_row[key] = float(value) else: processed_row[key] = value processed_data.append(processed_row) df = pl.DataFrame(processed_data)
3. 用Lazy模式加载
LazyFrame的类型推断逻辑更灵活,先转为LazyFrame再收集成DataFrame:
df = pl.LazyFrame(res).collect()
4. 排查隐藏的类型异常
有时候表面类型一致,实际存在None、numpy数值等特殊值,遍历数据排查:
for idx, row in enumerate(res): for col, val in row.items(): print(f"行{idx},列{col}:值={val},类型={type(val)}")
找到异常值后针对性处理即可。
内容的提问来源于stack exchange,提问作者Mohammad Talaei
相关产品推荐
相关产品推荐

