如何在Pandas中按特定行条件生成符合要求的新列?
Pandas生成符合规则的desired_output列
原始DataFrame
| exit | new_column |
|---|---|
| 0 | 0 |
| 0 | 0 |
| 1 | 1 |
| 0 | 0 |
| 0 | 0 |
| 0 | 0 |
| 0 | 0 |
| 0 | 0 |
| 1 | 1 |
生成规则
- 若exit列值为0,desired_output列设为0;
- 若exit列值为1,检查后续10行内是否存在另一个1:
- 存在:当前行desired_output设为1,后续第一个1所在行设为0;
- 不存在:当前行desired_output设为1。
问题代码
用户尝试了以下代码,但未得到预期结果:
df['new_column'] = 0 for i, row in combined_df.iterrows(): if row['exit'] == 1: next_rows = df.loc[i+1:i+10, 'exit'] if (next_rows == 1).any(): df.loc[i, 'new_column'] = 1 later_occurrence_index = next_rows[next_rows == 1].index[0] df.loc[later_occurrence_index, 'new_column'] = 0 else: df.loc[i, 'new_column'] = 1 else: df.loc[i, 'new_column'] = 0
代码问题分析
- 变量名不一致:循环遍历
combined_df,但操作的是df,可能导致数据不匹配; - 重复处理:当后续的1被标记为0后,循环到该行时,仍会执行检查逻辑,可能覆盖之前的标记;
iterrows()效率低下,且容易引发索引相关的错误。
正确解决方案
先提取所有exit=1的索引位置,逐个处理并跳过已标记的行:
import pandas as pd # 构造示例DataFrame data = {'exit': [0,0,1,0,0,0,0,0,1]} df = pd.DataFrame(data) df['desired_output'] = 0 # 获取所有exit=1的索引列表 exit_indices = df[df['exit'] == 1].index.tolist() for idx in exit_indices: # 若当前行已被标记为0,直接跳过 if df.loc[idx, 'desired_output'] == 0: continue # 检查后续10行范围内的exit值 next_range = df.loc[idx+1 : idx+10, 'exit'] # 找到后续第一个1的索引 next_one_idx = next_range[next_range == 1].index if not next_one_idx.empty: df.loc[idx, 'desired_output'] = 1 # 标记后续第一个1的行为0 df.loc[next_one_idx[0], 'desired_output'] = 0 else: df.loc[idx, 'desired_output'] = 1 print(df)
运行后输出结果:
exit desired_output 0 0 0 1 0 0 2 1 1 3 0 0 4 0 0 5 0 0 6 0 0 7 0 0 8 1 0
该结果完全符合规则:第2行的exit=1,后续10行内存在第8行的1,因此第2行设为1、第8行设为0;其余exit=0的行均为0。
内容的提问来源于stack exchange,提问作者deeni
相关产品推荐
相关产品推荐

