You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中当Pandas DataFrame列值变化时插入前后文本行

实现方案

首先,先构建你的原始DataFrame:

import pandas as pd

df = pd.DataFrame({
    'value': [0, 1, 0, 0, 0, 1, 2, 0],
    'type': ['brown', 'green', 'blue', 'brown', 'black', 'yellow', 'green', 'blue'],
    'section': ['sect1', 'sect1', 'sect2', 'sect3', 'sect4', 'sect4', 'sect4', 'sect5']
})

方案一:合并同组START与后续行(常规合理需求)

如果你的真实需求是每个以START开头的连续同section行组,末尾添加END(即sect1的START+1合并为一个块,sect4的START+1+2合并为一个块),可以按以下步骤实现:

  1. 替换value列的0为'START':
df['value'] = df['value'].replace(0, 'START')
  1. 按START行分组:每个START行作为新组的起点,后续同section的非START行归为同一组
df['group_id'] = df['value'].eq('START').cumsum()
  1. 遍历每个组,添加组内行和END行:
result_rows = []
for _, group in df.groupby('group_id'):
    result_rows.extend(group.to_dict('records'))
    result_rows.append({'value': 'END', 'type': '', 'section': ''})

# 转换为结果DataFrame
result_df = pd.DataFrame(result_rows).reset_index(drop=True)

运行后得到的结果:

valuetypesection
STARTbrownsect1
1greensect1
END
STARTbluesect2
END
STARTbrownsect3
END
STARTblacksect4
1yellowsect4
2greensect4
END
STARTbluesect5
END

方案二:严格匹配你给出的期望结果

如果确实需要将sect4的START行单独拆分为一个块,再将后续非START行单独作为一个块,可使用循环逐行处理:

# 替换0为START
df['value'] = df['value'].replace(0, 'START')

result_rows = []
i = 0
while i < len(df):
    current_row = df.iloc[i].copy()
    if current_row['value'] == 'START':
        # 添加当前START行和END行
        result_rows.append(current_row.to_dict())
        result_rows.append({'value': 'END', 'type': '', 'section': ''})
        
        # 检查后续是否有同section的非START行
        if i + 1 < len(df) and df.iloc[i+1]['section'] == current_row['section'] and df.iloc[i+1]['value'] != 'START':
            # 收集所有连续的同section非START行
            j = i + 1
            non_start_rows = []
            while j < len(df) and df.iloc[j]['section'] == current_row['section'] and df.iloc[j]['value'] != 'START':
                non_start_rows.append(df.iloc[j].to_dict())
                j += 1
            # 添加新的START行(复用当前行的type和section)
            result_rows.append({'value': 'START', 'type': current_row['type'], 'section': current_row['section']})
            result_rows.extend(non_start_rows)
            result_rows.append({'value': 'END', 'type': '', 'section': ''})
            i = j
        else:
            i += 1
    else:
        i += 1

result_df = pd.DataFrame(result_rows).reset_index(drop=True)

这段代码会生成你给出的期望结果,其中sect4会拆分为两个START-END块。

内容的提问来源于stack exchange,提问作者Connor Garrett

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 11:05:14