You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效替换Pandas DataFrame每行中1.0与0.0为指定内容?

高效转换Pandas DataFrame的0/1值为指定内容

需求说明

处理Pandas DataFrame时需实现以下转换:

  • 将每行中的0.0替换为空字符串('')
  • 将每行中的1.0替换为该行的索引值
    注:所有数据值仅为1.0或0.0

示例数据

初始DataFrame代码

import pandas as pd
# starting df
df = pd.DataFrame.from_dict({'A':[1.0,0.0,0.0],'B':[1.0,1.0,0.0],'C':[0.0,1.0,1.0]})
df.index=['x','y','z']
print(df)

输入样式

A    B    C
x  1.0  1.0  0.0
y  0.0  1.0  1.0
z  0.0  0.0  1.0

期望输出样式

A  B  C
x  x  x   
y     y  y
z        z

当前低效实现

原代码通过逐行循环处理,在(4548, 2044)规模的数据下效率极低:

for idx in df.index:
    df.loc[idx] = df.loc[idx].map(str).replace('1.0',str(idx))
    df.loc[idx] = df.loc[idx].map(str).replace('0.0','')

高效实现方案

利用Pandas/Numpy的矢量化操作替代循环,可大幅提升处理速度:

方案1:Pandas广播+where方法

# 生成与原DataFrame形状一致的索引值矩阵
idx_matrix = pd.DataFrame([df.index]*len(df.columns)).T
# 1.0位置保留索引值,0.0位置替换为空字符串
result_df = idx_matrix.where(df == 1.0, '')

方案2:applymap逐元素处理(比循环高效)

result_df = df.apply(lambda row: row.map(lambda x: row.name if x == 1.0 else ''), axis=1)

方案3:Numpy向量化操作(性能最优)

借助Numpy广播特性实现最快处理:

import numpy as np

# 将索引广播为与原DataFrame同形状的数组
idx_arr = np.tile(df.index.values.reshape(-1,1), (1, df.shape[1]))
# 批量替换值后转回DataFrame
result_df = pd.DataFrame(
    np.where(df.values == 1.0, idx_arr, ''),
    index=df.index,
    columns=df.columns
)

性能说明

针对(4548, 2044)规模的数据,矢量化操作的处理速度会比原循环快数倍至数十倍,彻底规避逐行处理的性能开销。

内容的提问来源于stack exchange,提问作者frustrated_bioinformatician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 16:40:19