You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于条件向Pandas DataFrame添加并转换行?

基于条件拆分DataFrame行的实现方案

问题描述

现有包含日期与条件的DataFrame,需根据Condition字段的值拆分行:当Condition=1时,将该行拆分为两行,分别使用Start作为起止日期、End作为起止日期;当Condition=0时,保留原行不变。

原始DataFrame

Start       End         Condition
03.10.2022  03.10.2022  0
03.10.2022  04.10.2022  1
03.10.2022  03.10.2022  0

预期结果

Start       End         
03.10.2022  03.10.2022  
03.10.2022  03.10.2022
04.10.2022  04.10.2022  
03.10.2022  03.10.2022  

尝试的错误代码

原计划用pd.explode实现,但以下代码报形状错误:

df["new_col"] = np.where(df['Condition'] == 1, 
                         df[['Start', 'End']].values.tolist(),
                         df['Start'])

解决方案

问题出在np.where的返回值形状不匹配:当Condition=1时返回的是每行对应一个包含两个元素的列表,但Condition=0时返回的是单个字符串,两者结构不一致。

正确的做法是让两种情况都返回列表格式,列表中的每个元素是(Start, End)格式的元组,这样explode后才能正确拆分并对应列。

完整代码

import pandas as pd
import numpy as np

# 构造原始DataFrame
df = pd.DataFrame({
    'Start': ['03.10.2022', '03.10.2022', '03.10.2022'],
    'End': ['03.10.2022', '04.10.2022', '03.10.2022'],
    'Condition': [0, 1, 0]
})

# 生成待拆分的列表列:每个元素是(Start, End)的元组
df['rows'] = np.where(
    df['Condition'] == 1,
    # Condition=1时,生成两个元组:(Start, Start)和(End, End)
    df.apply(lambda x: [(x['Start'], x['Start']), (x['End'], x['End'])], axis=1),
    # Condition=0时,生成单个元组:(Start, End)
    df.apply(lambda x: [(x['Start'], x['End'])], axis=1)
)

# 拆分列表列,然后扩展成Start和End列
result = df.explode('rows')[['rows']].apply(pd.Series, index=['Start', 'End'])

print(result)

代码解释

  • 用np.where根据条件生成rows列:
    • 当Condition=1,通过apply生成包含两个元组的列表,每个元组对应拆分后的一行起止日期
    • 当Condition=0,生成包含单个原行起止日期元组的列表
  • 使用explode将列表拆分成多行
  • 把rows列的元组扩展成Start和End两列,得到最终结果

内容的提问来源于stack exchange,提问作者Stacky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 17:50:35