Python中基于条件合并行:将Title列nan行合并至最近非nan行
解决方案:合并pandas DataFrame中Title为"nan"行的Content到最近有效行
实现函数
以下函数可高效处理大规模数据集,将Title列值为"nan"的行的Content合并到最近的非"nan"Title行,并删除原"nan"行:
import pandas as pd def merge_nan_title_content(df): # 创建分组键:为每个有效Title及其后续的nan行分配同一组编号 df['group_key'] = (df['Title'] != 'nan').cumsum() # 分组聚合处理 merged_df = df.groupby('group_key').agg( Title=('Title', 'first'), # 保留组内第一个有效Title Content=('Content', lambda x: ''.join(x)), # 拼接组内所有Content Measure=('Measure', 'first') # 保留组内第一个Measure值(同组Measure通常一致) ).reset_index(drop=True) return merged_df
函数逻辑说明
- 分组键生成:通过
(df['Title'] != 'nan').cumsum()生成连续的分组编号,每遇到一个非"nan"的Title,分组号递增,确保后续所有"nan"行与最近的有效Title行归为同一组。 - 聚合操作:
Title取每组第一个值,即该组的有效标题;Content将组内所有行的内容直接拼接,若需要分隔符(如空格、逗号),可将''.join(x)改为' '.join(x)或','.join(x);Measure取每组第一个值,若真实数据中同组Measure存在差异,可根据需求调整为'unique'等聚合方式。
验证示例
使用你提供的原始数据测试函数:
# 生成原始表格 data = {'Title': ['Background', 'Method', 'nan', 'Background', 'Method', 'Background', 'Method', 'nan', 'nan'], 'Content': ['text1', 'abc', 'dfg', 'text2', 'abcdfg', 'text3', 'ab', 'cd', 'fg'], 'Measure': ['Measure1', 'Measure1', 'Measure1', 'Measure2', 'Measure2', 'Measure3', 'Measure3', 'Measure3', 'Measure3']} df = pd.DataFrame(data) # 执行处理 processed_df = merge_nan_title_content(df) print(processed_df)
输出结果与目标表格一致:
Title Content Measure 0 Background text1 Measure1 1 Method abcdfg Measure1 2 Background text2 Measure2 3 Method abcdfg Measure2 4 Background text3 Measure3 5 Method abcd fg Measure3
内容的提问来源于stack exchange,提问作者Mando
相关产品推荐
相关产品推荐

