You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars DataFrame中并行处理过滤表达式索引生成稀疏矩阵?

优化Polars中生成稀疏矩阵的并行处理性能问题

问题描述

需要在Polars中根据一组过滤表达式获取匹配的行索引,据此生成稀疏矩阵(匹配位置为1,否则为0)。当前采用循环逐个过滤表达式的暴力实现,存在严重性能瓶颈。

现有低效代码

def get_sparse_matrix(exprs: list[pl.Expr]) -> scipy.sparse.csc_matrix:
    df = df.with_row_index('_index')
    rows: list[int] = []
    cols: list[int] = []
    for col, expr in enumerate(exprs):
        r = self.df.filter(expr)['_index']
        rows.extend(r)
        cols.extend([col] * len(r))

    X = csc_matrix((np.ones(len(rows)), (rows, cols)), shape= 
   (len(self.df), len(rules)))

    return X

示例输入

# 8行5列的Polars DataFrame(补充column_4以满足表达式需求)
df = pl.DataFrame({
    "column_0": [1,2,3,4,5,6,7,8],
    "column_1": [3,4,5,6,7,8,9,10],
    "column_2": [5,6,7,8,9,10,11,12],
    "column_3": [5,6,41,8,21,10,51,12],
    "column_4": [0,0,0,0,22,0,0,0]
})

# 三个过滤表达式
exprs = [pl.col('column_0') > 3, pl.col('column_1') < 6, pl.col('column_4') > 11]

示例输出

生成大小为8(记录数)×3(表达式数)的稀疏矩阵,当第i条记录匹配第j个表达式时,矩阵(i,j)位置元素为1。转换为密集矩阵后如下:

[[0 1 0]
 [0 1 0]
 [0 1 0]
 [1 0 0]
 [1 0 1]
 [1 0 0]
 [1 0 0]
 [1 0 0]]

并行优化方案

核心优化思路

  1. 批量计算布尔结果:一次性计算所有表达式的匹配结果,避免多次遍历DataFrame,Polars默认会并行处理多列计算。
  2. 长格式转换提取匹配对:用melt将布尔矩阵转为长格式,筛选出匹配的行,直接获取行索引和表达式列的对应关系。
  3. 减少Python循环开销:直接从Polars Series转换为NumPy数组,避免列表extend等低效操作。

优化后代码

import polars as pl
import numpy as np
from scipy.sparse import csc_matrix

def get_sparse_matrix(df: pl.DataFrame, exprs: list[pl.Expr]) -> csc_matrix:
    # 批量计算所有表达式的布尔匹配结果,每列对应一个表达式
    bool_results = df.select(exprs)
    # 添加行索引(若已有唯一索引列,可替换为原索引)
    bool_results = bool_results.with_row_index("_row_idx")
    # 转换为长格式,筛选出匹配成功的记录
    matched_pairs = bool_results.melt(
        id_vars=["_row_idx"],
        value_name="is_matched"
    ).filter(pl.col("is_matched"))
    
    # 提取行、列索引的NumPy数组
    rows = matched_pairs["_row_idx"].to_numpy()
    cols = matched_pairs["variable"].cast(pl.UInt32).to_numpy()
    
    # 构造CSC稀疏矩阵
    return csc_matrix(
        (np.ones(len(rows), dtype=np.int8), (rows, cols)),
        shape=(len(df), len(exprs))
    )

优化效果说明

  • 计算效率:Polars会并行处理所有表达式的计算,相比循环逐个过滤,遍历DataFrame的次数从N个表达式减少到1次,大幅降低IO和计算开销。
  • 内存效率:直接使用Polars的向量化操作和NumPy数组,避免Python列表的内存拷贝和循环开销。
  • 可扩展性:表达式数量越多,优化效果越明显,适合处理大规模数据集和大量表达式的场景。

内容的提问来源于stack exchange,提问作者xxx222

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 06:50:28