单行DataFrame与多行DataFrame逐元素比较,如何筛选符合条件的列号?
我来帮你解决这个问题——你的思路方向是对的,但问题出在np.where的返回值结构和apply的处理方式上,咱们一步步来修正:
为什么你的代码会得到空DataFrame?
当你执行df.apply(lambda x : np.where(x < thresh_df.values))时:
thresh_df.values是二维数组(形状(1,200)),而x是一维Series,虽然比较能正常触发广播,但np.where返回的是一个包含两个数组的元组:第一个是布尔矩阵中True的行索引(在x的上下文里全是0,因为x是单行),第二个才是我们需要的列索引。apply会把每个lambda的返回值(也就是这个元组)当成新行的元素,但元组的结构和DataFrame的行列匹配逻辑冲突,最终导致生成的结果看起来是空的。
正确的解决方案
我们直接利用pandas/numpy的矢量化特性,高效获取每行中小于对应阈值的列号,这里提供两种实用方法:
方法1:基于pandas的mask筛选(易读性高)
import pandas as pd import numpy as np # 先把阈值转成一维数组/Series,方便广播匹配 thresh_series = thresh_df.iloc[0] # 生成布尔矩阵:每个元素标记是否小于对应列的阈值 mask = df < thresh_series # 对每行筛选出值为True的列名(或列号),转成列表 result = mask.apply(lambda row: row[row].index.tolist(), axis=1)
result会是一个Series,每个元素对应df中一行的符合条件的列名列表;如果需要列号(整数索引),把row[row].index换成row[row].index.to_list()或者直接用np.where(row)[0].tolist()。
方法2:基于numpy的矢量化操作(性能更优,适合大数据集)
如果你的df行数很多(比如1000行以上),用numpy的原生操作会更快:
# 把阈值转成一维数组 thresh_array = thresh_df.values.flatten() # 获取所有符合条件的元素的行、列索引 row_indices, col_indices = np.where(df.values < thresh_array) # 按行分组,整理出每行对应的列号 result = pd.Series([col_indices[row_indices == i].tolist() for i in range(df.shape[0])])
测试示例
用小数据验证效果:
# 构造测试数据 df = pd.DataFrame([[1,3,5], [2,4,6], [0,2,4]], columns=['A','B','C']) thresh_df = pd.DataFrame([[2,3,5]], columns=['A','B','C']) # 用方法1执行 thresh_series = thresh_df.iloc[0] mask = df < thresh_series result = mask.apply(lambda row: row[row].index.tolist(), axis=1) print(result) # 输出: # 0 [A] # 1 [A, B] # 2 [A, B, C] # dtype: object
内容的提问来源于stack exchange,提问作者Fasty
相关产品推荐
相关产品推荐

