Polars read_excel指定列遇异常:文档与实际不符该如何解决?
Polars read_excel 指定列加载问题与解决方法
问题背景
根据Polars官方文档描述,read_excel的columns参数可传入列名或索引序列,因此编写了以下测试代码:
import polars df = polars.read_excel( "/Volumes/Spare/foo.xlsx", engine="calamine", sheet_name="natsav", read_options={"header_row": 2}, columns=(1,2,4,5,6,7), # 不需要第0和第3列 ) print(df.head())
但运行时抛出异常:
_fastexcel.InvalidParametersError: invalid parameters: `use_columns` callable could not be called (TypeError: 'tuple' object is not callable)
后续尝试使用返回布尔值的可调用对象:
def colspec(c): print(type(c)) return True
将columns设为colspec后程序正常运行,发现传入的参数类型为builtins.ColumnInfoNoDtype,但该类型无官方文档说明。由此产生两个疑问:文档是否存在错误?如何正确使用polars.read_excel加载指定列?
问题原因与解决方法
Polars官方文档对columns参数的描述存在滞后性——当使用calamine引擎时,columns参数仅支持可调用对象或列名字符串列表,不支持索引序列。这是因为calamine引擎底层的参数映射逻辑与Polars通用规则不同,导致索引序列被误判为可调用对象,从而抛出错误。
以下是几种正确加载指定列的方式:
1. 使用列名字符串列表(已知列名时)
如果提前知道目标列的名称,可以直接传入列名列表:
import polars df = polars.read_excel( "/Volumes/Spare/foo.xlsx", engine="calamine", sheet_name="natsav", read_options={"header_row": 2}, columns=["列名1", "列名2", "列名4"] # 替换为实际列名 )
2. 使用可调用对象基于列索引筛选
通过可调用对象接收ColumnInfoNoDtype参数,利用其index属性判断是否保留该列:
import polars def select_columns(col_info): # 保留索引为1、2、4、5、6、7的列 return col_info.index in {1,2,4,5,6,7} df = polars.read_excel( "/Volumes/Spare/foo.xlsx", engine="calamine", sheet_name="natsav", read_options={"header_row": 2}, columns=select_columns )
3. 先读取全量数据再筛选(适合小文件)
如果Excel文件体积不大,可以先读取所有列,再通过select方法选择目标列:
import polars df = polars.read_excel( "/Volumes/Spare/foo.xlsx", engine="calamine", sheet_name="natsav", read_options={"header_row": 2}, ).select([1,2,4,5,6,7]) # 按索引选择列
内容的提问来源于stack exchange,提问作者jackal
相关产品推荐
相关产品推荐

