You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历DataFrame替换0值为相邻列均值?代码调试求助

解决DataFrame中0值替换为同行相邻列均值的问题

问题说明

需求:遍历DataFrame所有元素,当检测到值为0时,将其替换为同一行相邻左右两列的均值。
用户提供的代码无法正常运行,期望将示例中第2行第2列的0替换为1和-2的均值(即-0.5)。

用户原代码:

import pandas as pd
import numpy as np
import numpy as np

d = {'col1': [4, 1, 2], 'col2': [1, 0, 4], 'col3': [12, -2, 4]}
df = pd.DataFrame(data=d)


for i in range(len(df)):
        for j in list(df):
                if df.loc[i,j] is np.nan:
                        df.loc[i,j] = (df.loc[i,j-1] + df.loc[i, j+1])/2

代码问题分析

  • 重复导入numpy:代码中两次执行import numpy as np,属于冗余操作
  • 判断条件错误:需求是检测0,但代码中判断的是np.nan,完全不符合需求
  • 列名处理错误:j是列名字符串(如'col1'),使用j-1或j+1会触发字符串加减的错误,无法获取相邻列
  • 循环方式低效:嵌套循环遍历DataFrame元素不符合pandas的最佳实践,大数据量下性能较差

解决方案

方案1:修正循环逻辑(基础版)

通过列索引位置来定位相邻列,修正判断条件:

import pandas as pd
import numpy as np

d = {'col1': [4, 1, 2], 'col2': [1, 0, 4], 'col3': [12, -2, 4]}
df = pd.DataFrame(data=d)

# 建立列名与索引的映射
col_list = df.columns.tolist()
col_index_map = {col: idx for idx, col in enumerate(col_list)}

for row_idx in range(len(df)):
    for col_name in col_list:
        # 检测当前元素是否为0
        if df.loc[row_idx, col_name] == 0:
            current_col_idx = col_index_map[col_name]
            # 确保当前列不是首尾列(首尾列没有左右相邻列)
            if 0 < current_col_idx < len(col_list) - 1:
                left_col = col_list[current_col_idx - 1]
                right_col = col_list[current_col_idx + 1]
                # 计算相邻列均值并替换
                df.loc[row_idx, col_name] = (df.loc[row_idx, left_col] + df.loc[row_idx, right_col]) / 2

print(df)

执行后输出:

col1  col2  col3
0     4   1.0    12
1     1  -0.5    -2
2     2   4.0     4

方案2:向量化操作(高效版)

利用pandas的shift方法实现向量化计算,避免循环,性能更优:

import pandas as pd
import numpy as np

d = {'col1': [4, 1, 2], 'col2': [1, 0, 4], 'col3': [12, -2, 4]}
df = pd.DataFrame(data=d)

# 计算每行相邻列的均值:shift(axis=1)左移一列,shift(-1,axis=1)右移一列
neighbor_mean = (df.shift(axis=1) + df.shift(-1, axis=1)) / 2
# 将原DataFrame中的0值替换为对应的相邻列均值,其余值保持不变
df = df.where(df != 0, neighbor_mean)

print(df)

输出结果与方案1一致,该方法适合处理大规模数据集。

内容的提问来源于stack exchange,提问作者Davi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 11:55:26