Python中当Pandas DataFrame列值变化时插入前后文本行
实现方案
首先,先构建你的原始DataFrame:
import pandas as pd df = pd.DataFrame({ 'value': [0, 1, 0, 0, 0, 1, 2, 0], 'type': ['brown', 'green', 'blue', 'brown', 'black', 'yellow', 'green', 'blue'], 'section': ['sect1', 'sect1', 'sect2', 'sect3', 'sect4', 'sect4', 'sect4', 'sect5'] })
方案一:合并同组START与后续行(常规合理需求)
如果你的真实需求是每个以START开头的连续同section行组,末尾添加END(即sect1的START+1合并为一个块,sect4的START+1+2合并为一个块),可以按以下步骤实现:
- 替换
value列的0为'START':
df['value'] = df['value'].replace(0, 'START')
- 按START行分组:每个START行作为新组的起点,后续同section的非START行归为同一组
df['group_id'] = df['value'].eq('START').cumsum()
- 遍历每个组,添加组内行和END行:
result_rows = [] for _, group in df.groupby('group_id'): result_rows.extend(group.to_dict('records')) result_rows.append({'value': 'END', 'type': '', 'section': ''}) # 转换为结果DataFrame result_df = pd.DataFrame(result_rows).reset_index(drop=True)
运行后得到的结果:
| value | type | section |
|---|---|---|
| START | brown | sect1 |
| 1 | green | sect1 |
| END | ||
| START | blue | sect2 |
| END | ||
| START | brown | sect3 |
| END | ||
| START | black | sect4 |
| 1 | yellow | sect4 |
| 2 | green | sect4 |
| END | ||
| START | blue | sect5 |
| END |
方案二:严格匹配你给出的期望结果
如果确实需要将sect4的START行单独拆分为一个块,再将后续非START行单独作为一个块,可使用循环逐行处理:
# 替换0为START df['value'] = df['value'].replace(0, 'START') result_rows = [] i = 0 while i < len(df): current_row = df.iloc[i].copy() if current_row['value'] == 'START': # 添加当前START行和END行 result_rows.append(current_row.to_dict()) result_rows.append({'value': 'END', 'type': '', 'section': ''}) # 检查后续是否有同section的非START行 if i + 1 < len(df) and df.iloc[i+1]['section'] == current_row['section'] and df.iloc[i+1]['value'] != 'START': # 收集所有连续的同section非START行 j = i + 1 non_start_rows = [] while j < len(df) and df.iloc[j]['section'] == current_row['section'] and df.iloc[j]['value'] != 'START': non_start_rows.append(df.iloc[j].to_dict()) j += 1 # 添加新的START行(复用当前行的type和section) result_rows.append({'value': 'START', 'type': current_row['type'], 'section': current_row['section']}) result_rows.extend(non_start_rows) result_rows.append({'value': 'END', 'type': '', 'section': ''}) i = j else: i += 1 else: i += 1 result_df = pd.DataFrame(result_rows).reset_index(drop=True)
这段代码会生成你给出的期望结果,其中sect4会拆分为两个START-END块。
内容的提问来源于stack exchange,提问作者Connor Garrett
相关产品推荐
相关产品推荐

