You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars read_excel指定列遇异常:文档与实际不符该如何解决?

Polars read_excel 指定列加载问题与解决方法

问题背景

根据Polars官方文档描述,read_excel的columns参数可传入列名或索引序列,因此编写了以下测试代码:

import polars

df = polars.read_excel(
    "/Volumes/Spare/foo.xlsx",
    engine="calamine",
    sheet_name="natsav",
    read_options={"header_row": 2},
    columns=(1,2,4,5,6,7), # 不需要第0和第3列
)

print(df.head())

但运行时抛出异常:

_fastexcel.InvalidParametersError: invalid parameters: `use_columns` callable could not be called (TypeError: 'tuple' object is not callable)

后续尝试使用返回布尔值的可调用对象:

def colspec(c):
    print(type(c))
    return True

将columns设为colspec后程序正常运行,发现传入的参数类型为builtins.ColumnInfoNoDtype,但该类型无官方文档说明。由此产生两个疑问:文档是否存在错误?如何正确使用polars.read_excel加载指定列?

问题原因与解决方法

Polars官方文档对columns参数的描述存在滞后性——当使用calamine引擎时,columns参数仅支持可调用对象或列名字符串列表,不支持索引序列。这是因为calamine引擎底层的参数映射逻辑与Polars通用规则不同,导致索引序列被误判为可调用对象,从而抛出错误。

以下是几种正确加载指定列的方式:

1. 使用列名字符串列表(已知列名时)

如果提前知道目标列的名称,可以直接传入列名列表:

import polars

df = polars.read_excel(
    "/Volumes/Spare/foo.xlsx",
    engine="calamine",
    sheet_name="natsav",
    read_options={"header_row": 2},
    columns=["列名1", "列名2", "列名4"] # 替换为实际列名
)

2. 使用可调用对象基于列索引筛选

通过可调用对象接收ColumnInfoNoDtype参数,利用其index属性判断是否保留该列:

import polars

def select_columns(col_info):
    # 保留索引为1、2、4、5、6、7的列
    return col_info.index in {1,2,4,5,6,7}

df = polars.read_excel(
    "/Volumes/Spare/foo.xlsx",
    engine="calamine",
    sheet_name="natsav",
    read_options={"header_row": 2},
    columns=select_columns
)

3. 先读取全量数据再筛选(适合小文件)

如果Excel文件体积不大,可以先读取所有列,再通过select方法选择目标列:

import polars

df = polars.read_excel(
    "/Volumes/Spare/foo.xlsx",
    engine="calamine",
    sheet_name="natsav",
    read_options={"header_row": 2},
).select([1,2,4,5,6,7]) # 按索引选择列

内容的提问来源于stack exchange,提问作者jackal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 21:59:56