You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单行DataFrame与多行DataFrame逐元素比较,如何筛选符合条件的列号?

我来帮你解决这个问题——你的思路方向是对的,但问题出在np.where的返回值结构和apply的处理方式上,咱们一步步来修正:

为什么你的代码会得到空DataFrame?

当你执行df.apply(lambda x : np.where(x < thresh_df.values))时:

  • thresh_df.values是二维数组(形状(1,200)),而x是一维Series,虽然比较能正常触发广播,但np.where返回的是一个包含两个数组的元组:第一个是布尔矩阵中True的行索引(在x的上下文里全是0,因为x是单行),第二个才是我们需要的列索引。
  • apply会把每个lambda的返回值(也就是这个元组)当成新行的元素,但元组的结构和DataFrame的行列匹配逻辑冲突,最终导致生成的结果看起来是空的。

正确的解决方案

我们直接利用pandas/numpy的矢量化特性,高效获取每行中小于对应阈值的列号,这里提供两种实用方法:

方法1:基于pandas的mask筛选(易读性高)

import pandas as pd
import numpy as np

# 先把阈值转成一维数组/Series,方便广播匹配
thresh_series = thresh_df.iloc[0]

# 生成布尔矩阵:每个元素标记是否小于对应列的阈值
mask = df < thresh_series

# 对每行筛选出值为True的列名(或列号),转成列表
result = mask.apply(lambda row: row[row].index.tolist(), axis=1)

result会是一个Series,每个元素对应df中一行的符合条件的列名列表;如果需要列号(整数索引),把row[row].index换成row[row].index.to_list()或者直接用np.where(row)[0].tolist()。

方法2:基于numpy的矢量化操作(性能更优,适合大数据集)

如果你的df行数很多(比如1000行以上),用numpy的原生操作会更快:

# 把阈值转成一维数组
thresh_array = thresh_df.values.flatten()

# 获取所有符合条件的元素的行、列索引
row_indices, col_indices = np.where(df.values < thresh_array)

# 按行分组,整理出每行对应的列号
result = pd.Series([col_indices[row_indices == i].tolist() for i in range(df.shape[0])])

测试示例

用小数据验证效果:

# 构造测试数据
df = pd.DataFrame([[1,3,5], [2,4,6], [0,2,4]], columns=['A','B','C'])
thresh_df = pd.DataFrame([[2,3,5]], columns=['A','B','C'])

# 用方法1执行
thresh_series = thresh_df.iloc[0]
mask = df < thresh_series
result = mask.apply(lambda row: row[row].index.tolist(), axis=1)

print(result)
# 输出:
# 0        [A]
# 1     [A, B]
# 2    [A, B, C]
# dtype: object

内容的提问来源于stack exchange,提问作者Fasty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:57:52