You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量优化Pandas特征处理函数?求非并行化高效方案

高效批量处理多特征的实现方案

核心优化思路

原函数的瓶颈在于逐个特征重复执行DataFrame创建、排序、分位数计算等操作,我们通过批量矢量化操作替代单特征循环,减少重复开销,同时严格保留原函数的业务逻辑。

优化后的代码实现

import pandas as pd
import numpy as np

def process_all_columns(X, y, n_quantiles=20):
    # 将标签转换为numpy数组,提升索引效率
    y_np = y.to_numpy()
    # 生成分位数分割点
    quantile_points = np.linspace(0, 1, n_quantiles)
    
    # 1. 批量获取所有特征的排序索引(缺失值默认排最后)
    # 用numpy矢量化处理替代pandas apply,速度更快
    X_filled = X.fillna(np.inf)  # 将NaN替换为极大值,确保argsort时排到最后
    sorted_indices = np.argsort(X_filled.to_numpy(), axis=0)
    # 批量生成每个特征排序后的数值数组
    sorted_features = X.to_numpy()[sorted_indices, np.arange(X.shape[1])]
    
    # 2. 批量计算每个特征非缺失值的分位数
    thresholds = []
    for col in X.T:
        non_na_vals = col[~pd.isna(col)].to_numpy()
        thresholds.append(np.quantile(non_na_vals, quantile_points))
    thresholds = np.array(thresholds)  # shape: (n_features, n_quantiles)
    
    # 3. 批量计算每个特征的拆分索引
    indices = np.array([
        np.searchsorted(sorted_col, thresh, side='right')
        for sorted_col, thresh in zip(sorted_features.T, thresholds)
    ])  # shape: (n_features, n_quantiles)
    
    # 4. 组装结果(仅此处循环特征列,核心逻辑已批量完成)
    feature_results = {}
    for feat_idx, feature_name in enumerate(X.columns):
        # 获取当前特征对应的排序后标签
        sorted_y = y_np[sorted_indices[:, feat_idx]]
        # 生成拆分后的标签元组列表
        y_tuples = [
            (sorted_y[:i].tolist(), sorted_y[i:].tolist())
            for i in indices[feat_idx]
        ]
        feature_results[feature_name] = (y_tuples, thresholds[feat_idx])
    
    return feature_results

关键优化点说明

  • 避免重复DataFrame操作:不再为每个特征单独创建DataFrame,直接对整个X矩阵进行批量处理,大幅减少IO和对象创建开销。
  • 矢量化排序:用np.argsort一次性完成所有特征的排序索引计算,比pandasapply的循环排序效率更高,同时保留缺失值排最后的逻辑。
  • 批量分位数计算:通过一次循环完成所有特征的非缺失值分位数计算,结合numpy的矢量化运算加速。
  • 结果组装轻量化:仅在最后一步循环特征列组装结果,核心计算逻辑均已通过批量操作完成,循环次数等于特征数,而非原函数的「特征数×分位数数量」。

与原函数的一致性验证

  • 完全保留原函数的缺失值处理逻辑:排序时缺失值排最后,分位数计算基于非缺失值,拆分索引在全量特征列(含缺失)中查找。
  • 返回结果结构兼容:以字典形式返回,键为特征名,值为原函数返回的(y_tuples, thresholds)元组,可直接替换原循环调用逻辑。

内容的提问来源于stack exchange,提问作者Maxou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 04:50:14