如何在Polars中实现Pandas基于索引的查询并保持指定行顺序?
Polars 实现 Pandas.loc 指定顺序筛选的高效方案
问题背景
要实现 Pandas 中
df.loc[[23, 132]]的功能——按指定顺序筛选行,但 Polars 没有类似的索引机制。原 Pandas 代码输出行顺序与指定的[23,132]完全一致:import pandas as pd import polars as pl df = pd.DataFrame(data = ([21,123], [132,412], [23, 43]), columns = ['c1', 'c2']).set_index("c1") out = df.loc[[23, 132]] print(pl.from_pandas(out.reset_index()))但直接用 Polars 的
filter方法会按原数据顺序返回,无法匹配指定顺序:df = pl.DataFrame(data = ([21,123], [132,412], [23, 43]), schema = ['c1', 'c2'], orient = 'row') print(df.filter(pl.col("c1").is_in([23, 132])))输出行顺序为
[132,23],不符合需求。原数据有3000万行,需高效方案,避免排序操作。
高效解决方案
方案1:左连接临时表(最优,适配千万级数据)
利用 Polars 的左连接特性,创建包含指定顺序的临时表,通过连接保留目标顺序,性能高效且无需排序:
import polars as pl # 原数据 df = pl.DataFrame(data = ([21,123], [132,412], [23, 43]), schema = ['c1', 'c2'], orient = 'row') # 指定筛选顺序 target_values = [23, 132] # 创建带顺序标记的临时表 target_df = pl.DataFrame({"c1": target_values}).with_row_index("order_idx") # 左连接匹配数据,按临时表顺序整理后移除辅助列 result = target_df.join(df, on="c1", how="left").sort("order_idx").drop("order_idx") print(result)
输出结果:
shape: (2, 2) ┌─────┬─────┐ │ c1 ┆ c2 │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞═════╪═════╡ │ 23 ┆ 43 │ │ 132 ┆ 412 │ └─────┴─────┘
方案2:字典映射(适合小目标列表)
如果目标筛选列表较短,可以通过字典映射快速按顺序提取行,适合小批量场景:
import polars as pl df = pl.DataFrame(data = ([21,123], [132,412], [23, 43]), schema = ['c1', 'c2'], orient = 'row') target_values = [23, 132] # 构建c1到对应c2值的映射 value_map = df.select("c1", "c2").to_dict(as_series=False) c1_map = dict(zip(value_map['c1'], value_map['c2'])) # 按目标顺序提取有效行 filtered_data = [[val, c1_map[val]] for val in target_values if val in c1_map] result = pl.DataFrame(filtered_data, schema=['c1', 'c2']) print(result)
注意:若目标列表过长或原数据量极大,字典构建会占用较多内存,此时优先选择方案1。
内容的提问来源于stack exchange,提问作者RacGreen1
相关产品推荐
相关产品推荐

