如何基于指定索引合并DataFrame行并排除NaN值?
如何将含NaN的行与前一行合并(排除NaN值)
问题描述
现有如下数据集(实际包含更多列和行):
id type value 0 104 0 7999 1 105 1 196193579 2 108 0 245744 3 NaN 1 NaN
部分行存在NaN值,已获取这些行的索引(例如indexes=[3]),需要将这些行与前一行合并,合并时排除NaN值,最终得到的DataFrame如下:
id type value 0 104 0 7999 1 105 1 196193579 2 108 01 245744
注意:首行绝不会出现在索引列表中,解决方案需基于给定的索引列表,也可使用已知的含NaN的列名。
解决方案
实现思路
- 倒序遍历目标索引:避免删除行后索引偏移导致后续处理出错
- 逐列处理:对每个目标行,将其非NaN值与前一行的对应列值拼接(或替换前一行的NaN值)
- 删除处理后的目标行,最后重置索引(可选)
代码实现
基础版本
import pandas as pd import numpy as np # 示例数据集 df = pd.DataFrame({ 'id': [104, 105, 108, np.nan], 'type': [0, 1, 0, 1], 'value': [7999, 196193579, 245744, np.nan] }) def merge_nan_rows(df, indexes): # 倒序遍历索引,防止删除行后索引混乱 for idx in sorted(indexes, reverse=True): prev_idx = idx - 1 # 遍历所有列 for col in df.columns: curr_val = df.loc[idx, col] prev_val = df.loc[prev_idx, col] # 仅保留非NaN值,若两行都有值则拼接 if not pd.isna(curr_val): if not pd.isna(prev_val): df.loc[prev_idx, col] = str(prev_val) + str(curr_val) else: df.loc[prev_idx, col] = curr_val # 删除当前行 df = df.drop(idx) # 重置索引(按需选择) df = df.reset_index(drop=True) return df # 调用示例 indexes = [3] result_df = merge_nan_rows(df.copy(), indexes) print(result_df)
优化版本(指定含NaN的列)
如果已知存在NaN的列名,可以只处理这些列,提升效率:
def merge_nan_rows(df, indexes, nan_cols=None): # 只处理指定的含NaN列,默认处理所有列 cols_to_process = nan_cols if nan_cols is not None else df.columns for idx in sorted(indexes, reverse=True): prev_idx = idx - 1 for col in cols_to_process: curr_val = df.loc[idx, col] prev_val = df.loc[prev_idx, col] if not pd.isna(curr_val): if not pd.isna(prev_val): df.loc[prev_idx, col] = str(prev_val) + str(curr_val) else: df.loc[prev_idx, col] = curr_val df = df.drop(idx) df = df.reset_index(drop=True) return df # 调用示例(已知含NaN的列是id和value) indexes = [3] result_df = merge_nan_rows(df.copy(), indexes, nan_cols=['id', 'value'])
输出结果
运行上述代码后,得到的结果符合预期:
id type value 0 104 0 7999 1 105 1 196193579 2 108 01 245744
内容的提问来源于stack exchange,提问作者Learning from masters
相关产品推荐
相关产品推荐

