You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何重新推断已有Polars DataFrame的数据类型?

如何在Polars中重新推断过滤后DataFrame的Schema(避免逐个设置列类型)

我有一个CSV文件,部分行包含错误值(用字符串替代整数)。为了成功读取文件,必须设置infer_schema_length = 0,否则读取会失败,但这会把所有列都读成字符串类型。过滤掉错误行后,我想重新推断整个DataFrame的数据类型——由于列数很多,不想逐个手动设置类型,也无法编辑原CSV文件。

当前代码:

import polars as pl

ids_df = pl.read_csv(dataset_path, infer_schema_length=0)
filtered_df = ids_df.filter(~(pl.col("Label") == "Label"))
print(filtered_df.dtypes)
# 输出全是Utf8类型

解决方案

方法1:自动转换所有数值列

用pl.all()选中所有列,通过str.to_numeric()自动推断整数/浮点数类型,同时跳过无法转换的值:

# 自动转换所有列到合适的数值类型
converted_df = filtered_df.with_columns(
    pl.all().cast(pl.Utf8).str.to_numeric(downcast_integer=True, strict=False)
)

print(converted_df.dtypes)
# 现在会显示正确的数值类型(如Int64、Float64),无法转换的列仍保留Utf8
  • downcast_integer=True:能转成整数的列会自动用更紧凑的整数类型(比如Int32/Int64)
  • strict=False:遇到无法转换的值时不会报错,会转为null(需要严格校验的话设为True)

方法2:内存中转存后重新读取

把过滤后的数据转成内存中的CSV字符串,再让Polars重新读取并推断Schema:

# 把过滤后的数据写入内存CSV,再重新读取
converted_df = pl.read_csv(
    filtered_df.write_csv(),
    infer_schema_length=None  # 用全部数据推断Schema,默认是1000,设为None用全量数据
)

print(converted_df.dtypes)

这种方式和直接读取正常CSV的效果一致,Polars会自动根据所有行推断每列的最优类型。

方法3:针对性转换(如果部分列明确是字符串)

如果知道某些列肯定是字符串类型,可以只转换其他列:

# 比如除了"Label"列是字符串,其他都转数值类型
converted_df = filtered_df.with_columns(
    pl.exclude("Label").cast(pl.Utf8).str.to_numeric(downcast_integer=True, strict=False)
)

内容的提问来源于stack exchange,提问作者yarvis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 13:32:25