如何优化Polars中统计每行ref值在其他列出现次数的代码
Polars中用map_rows统计每行ref值在其他列的出现次数优化方案
我需要在Polars中使用map_rows方法添加一列,统计每行里ref列的值在其余列中出现的次数(排除ref列本身)。
示例数据
import polars as pl df = pl.from_repr(''' ┌─────┬─────┬─────┐ │ ref ┆ v1 ┆ v2 │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 │ ╞═════╪═════╪═════╡ │ -1 ┆ -1 ┆ -1 │ # -1 出现2次(排除ref) │ 2 ┆ 5 ┆ 5 │ # 2 出现0次 │ 8 ┆ 0 ┆ 8 │ # 8 出现1次 └─────┴─────┴─────┘ ''')
Pandas 实现代码
df_pd = df.to_pandas() df_pd['count'] = df_pd.apply(lambda r: len([i for i in r if i == r[0]]) - 1, axis=1) df_pd = df_pd.drop('ref', axis=1) df_pd
运行结果:
v1 v2 count 0 -1 -1 2 1 5 5 0 2 0 8 1
当前Polars实现(待优化)
虽然能得到预期结果,但rename和hstack的写法较为繁琐:
x = df.map_rows(lambda r: len([i for i in r if i == r[0]]) - 1).rename({'map': 'count'}) df = df.hstack([x.to_series()]).drop('ref') df
运行结果:
shape: (3, 3) ┌─────┬─────┬───────┐ │ v1 ┆ v2 ┆ count │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 │ ╞═════╪═════╪═══════╡ │ -1 ┆ -1 ┆ 2 │ │ 5 ┆ 5 ┆ 0 │ │ 0 ┆ 8 ┆ 1 │ └─────┴─────┴───────┘
优化方案
方案1:用with_columns直接添加列
利用with_columns链式调用,直接给map_rows的结果命名,省去单独处理中间结果的步骤:
df = df.with_columns( pl.map_rows( lambda r: len([i for i in r[1:] if i == r[0]]), # 直接统计ref之外的列,无需减1 return_dtype=pl.Int64 ).alias('count') ).drop('ref') print(df)
方案2:直接在map_rows中构造完整行数据
在map_rows的lambda里直接返回包含目标列的字典,一步生成最终DataFrame:
df = df.map_rows( lambda r: {'v1': r[1], 'v2': r[2], 'count': len([i for i in r[1:] if i == r[0]])} ).to_df() print(df)
两种方案都能得到预期结果,代码更紧凑简洁,避免了冗余的rename和hstack操作。
内容的提问来源于stack exchange,提问作者Quiescent
相关产品推荐
相关产品推荐

