You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Polars创建DataFrame遇ComputeError报错,Pandas可正常运行求解

Polars创建DataFrame时触发ComputeError的问题

错误信息

ComputeError: could not append value: 1.41431 of type: f64 to the builder; make sure that all rows have the same schema or consider increasing infer_schema_length it might also be that a value overflows the data-type's capacity

复现代码

import json
import pandas as pd
import polars as pl
import requests

# 加载测试JSON数据
response = requests.get('测试JSON文件地址')
res = json.loads(response.text)

# Pandas创建DataFrame正常运行
df_pd = pd.DataFrame(res)
# Polars创建DataFrame抛出上述错误
df = pl.DataFrame(res)

问题说明

原本预期pl.DataFrame(res)能和pd.DataFrame(res)一样正常执行,但Polars抛出了上述错误。使用的是最新版Polars(0.20.31),曾参考过Polars的一个已关闭议题,但问题仍未解决。

已尝试的操作均无效:

  • 检查过res中值的类型,表面看起来一致
  • 将浮点数保留5位小数
  • 增大infer_schema_length参数

解决思路

1. 强制指定Schema

Polars的类型推断比Pandas严格,哪怕某一列前N行是整数,后面出现浮点数就会触发错误。直接手动指定每列的数据类型,跳过自动推断:

# 根据实际数据结构定义schema,示例如下
schema = {
    "column1": pl.Float64,
    "column2": pl.String,
    # 其他列依次对应指定类型
}
df = pl.DataFrame(res, schema=schema)

2. 预处理统一数据类型

遍历所有行,把数值类型统一转为float64,避免混合类型导致推断失败:

processed_data = []
for row in res:
    processed_row = {}
    for key, value in row.items():
        # 把整数、浮点数统一转为float64
        if isinstance(value, (int, float)):
            processed_row[key] = float(value)
        else:
            processed_row[key] = value
    processed_data.append(processed_row)
df = pl.DataFrame(processed_data)

3. 用Lazy模式加载

LazyFrame的类型推断逻辑更灵活,先转为LazyFrame再收集成DataFrame:

df = pl.LazyFrame(res).collect()

4. 排查隐藏的类型异常

有时候表面类型一致,实际存在None、numpy数值等特殊值,遍历数据排查:

for idx, row in enumerate(res):
    for col, val in row.items():
        print(f"行{idx},列{col}:值={val},类型={type(val)}")

找到异常值后针对性处理即可。

内容的提问来源于stack exchange,提问作者Mohammad Talaei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 18:57:20