Pandas如何基于连续行空值条件合并对应文本列内容
实现方案
核心思路
通过判断连续行的Value是否为空的状态变化生成分组标识,连续空值行归为同一组,非空值行单独分组,之后按分组聚合即可完成Name字段拼接。
完整代码
import pandas as pd # 构造示例DataFrame dummy_df = pd.DataFrame([{'Name': 'First', 'Value': 1}, {'Name': 'Start of', 'Value': None}, {'Name': 'cut off', 'Value': None}, {'Name': 'Last', 'Value': 10}, {'Name': 'First of', 'Value': None}, {'Name': 'three lines', 'Value': None}, {'Name': 'cut off', 'Value': None}, {'Name': 'Actually last', 'Value': 100}]) # 生成分组标识:连续空值/非空值行会被分到同一组 is_null = dummy_df['Value'].isna() group_key = (is_null != is_null.shift()).cumsum() # 分组聚合:Name列用空格拼接,Value列取组内第一个值 result = dummy_df.groupby(group_key).agg({ 'Name': ' '.join, 'Value': 'first' }).reset_index(drop=True) print(result)
输出结果
运行后得到的结果和预期一致:
Name Value 0 First 1.0 1 Start of cut off NaN 2 Last 10.0 3 First of three lines cut off NaN 4 Actually last 100.0
内容的提问来源于stack exchange,提问作者TomNash
相关产品推荐
相关产品推荐

