You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Polars .filter切片比Pandas .loc慢,如何优化?

提升Polars切片操作速度的优化方案

在将Pandas代码迁移到Polars时,发现Polars的.filter操作在循环中处理单条日期切片时,速度远慢于Pandas的.loc。测试代码及结果如下:

import pandas as pd
import polars as pl
import datetime as dt
import numpy as np
import time

date_index = pd.date_range(dt.date(2001,1,1), dt.date(2020,1,1),freq='1H')
n = date_index.shape[0]
test_pd = pd.DataFrame(data = np.random.randint(1,100, n), index=date_index, columns = ['test'])
test_pl = pl.DataFrame(test_pd.reset_index())
test_dates = date_index[np.random.randint(0,n,1000)]

st = time.perf_counter()
for i in test_dates:
    d = test_pd.loc[i,:]
print(f"Pandas {time.perf_counter() - st}")


st = time.perf_counter()
for i in test_dates:
    d = test_pl.filter(index=i)
print(f"Polars {time.perf_counter() - st}")

# 输出结果
# Pandas 0.1854726000019582
# Polars 2.1125728000042727

核心优化思路

Polars的设计初衷是批量向量化操作,而非循环单条处理。原代码中循环调用.filter会重复触发查询解析、执行的开销,这是导致速度慢的根本原因。以下是两种高效优化方法:

方法1:给Polars DataFrame设置索引,使用.loc切片

和Pandas一样,给Polars的日期列设置为索引,利用索引的快速查找特性,直接用.loc实现和Pandas一致的高效切片:

# 给Polars DataFrame设置索引
test_pl_indexed = test_pl.set_index("index")

st = time.perf_counter()
for i in test_dates:
    d = test_pl_indexed.loc[i]
print(f"Polars with index {time.perf_counter() - st}")

该方法的速度会接近甚至超过Pandas的.loc,因为Polars的索引查找同样做了底层优化。

方法2:批量查询,避免循环

将所有待查询的日期一次性传入,用is_in实现批量过滤,彻底消除循环开销:

st = time.perf_counter()
# 批量过滤所有目标日期
result = test_pl.filter(pl.col("index").is_in(test_dates))
print(f"Polars batch filter {time.perf_counter() - st}")

这种方式的效率最高,因为Polars会一次性完成所有查询的向量化计算,完全避开循环的额外开销。如果需要后续处理单条数据,可以从批量结果中提取,整体效率仍远高于循环单条过滤。

效果说明

优化后的两种方法,速度都会大幅提升:

  • 设置索引后循环.loc的耗时会降到和Pandas同一量级(甚至更快)
  • 批量过滤的耗时会远低于循环操作(通常是毫秒级)

内容的提问来源于stack exchange,提问作者parmatma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 22:22:21