You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars:如何无报错选择非存在列,或筛选含指定列的数据集

解决方案

一、选择不存在的列时返回默认值/Null

要在Polars中选择可能不存在的列且不抛出异常,可以借助pl.col()的default参数,指定列不存在时的默认值(比如None对应Null)。示例代码如下:

import polars as pl

# 示例DataFrame
df = pl.DataFrame({"id": [1,2,3]})

# 选择存在的id列和不存在的bar列,bar列返回Null
result = df.select(
    pl.col("id"),
    pl.col("bar").default(None)
)
print(result)

输出结果:

shape: (3, 2)
┌─────┬──────┐
│ id  ┆ bar  │
│ --- ┆ ---  │
│ i64 ┆ null │
├─────┼──────┤
│ 1   ┆ null │
│ 2   ┆ null │
│ 3   ┆ null │
└─────┴──────┘

如果需要自定义默认值(比如空字符串),直接替换None即可:pl.col("bar").default("")。

二、仅合并包含指定列的CSV文件

针对你的场景,要只纳入包含bar列的CSV文件,需要先遍历匹配的文件,检查每个文件的表头,再筛选出符合条件的文件进行扫描合并。代码实现如下:

import polars as pl
import glob

# 获取所有匹配的CSV文件
csv_files = glob.glob("df*.csv")

# 筛选出包含"bar"列的文件
valid_files = []
for file in csv_files:
    # 仅读取表头,无需加载全量数据,提升效率
    df_schema = pl.read_csv(file, n_rows=0)
    if "bar" in df_schema.columns:
        valid_files.append(file)

# 扫描并合并符合条件的文件
df = pl.scan_csv(valid_files).select(["id", "bar"])
res = df.collect()
print(res)

最终结果仅包含df1.csv的内容:

shape: (3, 2)
┌─────┬───────┐
│ id  ┆ bar   │
│ --- ┆ ---   │
│ i64 ┆ str   │
├─────┼───────┤
│ 1   ┆ sugar │
│ 2   ┆ ham   │
│ 3   ┆ spam  │
└─────┴───────┘

内容的提问来源于stack exchange,提问作者lebesgue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 06:01:07