使用Polars .filter切片比Pandas .loc慢,如何优化?
提升Polars切片操作速度的优化方案
在将Pandas代码迁移到Polars时,发现Polars的.filter操作在循环中处理单条日期切片时,速度远慢于Pandas的.loc。测试代码及结果如下:
import pandas as pd import polars as pl import datetime as dt import numpy as np import time date_index = pd.date_range(dt.date(2001,1,1), dt.date(2020,1,1),freq='1H') n = date_index.shape[0] test_pd = pd.DataFrame(data = np.random.randint(1,100, n), index=date_index, columns = ['test']) test_pl = pl.DataFrame(test_pd.reset_index()) test_dates = date_index[np.random.randint(0,n,1000)] st = time.perf_counter() for i in test_dates: d = test_pd.loc[i,:] print(f"Pandas {time.perf_counter() - st}") st = time.perf_counter() for i in test_dates: d = test_pl.filter(index=i) print(f"Polars {time.perf_counter() - st}") # 输出结果 # Pandas 0.1854726000019582 # Polars 2.1125728000042727
核心优化思路
Polars的设计初衷是批量向量化操作,而非循环单条处理。原代码中循环调用.filter会重复触发查询解析、执行的开销,这是导致速度慢的根本原因。以下是两种高效优化方法:
方法1:给Polars DataFrame设置索引,使用.loc切片
和Pandas一样,给Polars的日期列设置为索引,利用索引的快速查找特性,直接用.loc实现和Pandas一致的高效切片:
# 给Polars DataFrame设置索引 test_pl_indexed = test_pl.set_index("index") st = time.perf_counter() for i in test_dates: d = test_pl_indexed.loc[i] print(f"Polars with index {time.perf_counter() - st}")
该方法的速度会接近甚至超过Pandas的.loc,因为Polars的索引查找同样做了底层优化。
方法2:批量查询,避免循环
将所有待查询的日期一次性传入,用is_in实现批量过滤,彻底消除循环开销:
st = time.perf_counter() # 批量过滤所有目标日期 result = test_pl.filter(pl.col("index").is_in(test_dates)) print(f"Polars batch filter {time.perf_counter() - st}")
这种方式的效率最高,因为Polars会一次性完成所有查询的向量化计算,完全避开循环的额外开销。如果需要后续处理单条数据,可以从批量结果中提取,整体效率仍远高于循环单条过滤。
效果说明
优化后的两种方法,速度都会大幅提升:
- 设置索引后循环
.loc的耗时会降到和Pandas同一量级(甚至更快) - 批量过滤的耗时会远低于循环操作(通常是毫秒级)
内容的提问来源于stack exchange,提问作者parmatma
相关产品推荐
相关产品推荐

