You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas apply处理Timestamp列时自定义函数返回None问题

问题根因

返回None是因为get_start_times()函数的条件分支没有覆盖所有可能的取值场景,当行数据同时不满足两个判断条件时,Python函数没有显式返回值,会默认返回None。

未被覆盖的具体场景:

  • 不满足第一个条件:x["site_time"].hour 大于6(即7~23点区间)
  • 同时不满足第二个条件:x["site_time"].hour 小于 x["shift_start_time"].hour

举个实际触发的例子:site_time为当日9点(hour=9,不满足0~6点的第一个条件),shift_start_time为当日10点(hour=10,此时shift_start_time.hour <= site_time.hour即10<=9不成立,第二个条件也不满足),两个分支都不会触发,函数直接返回默认值None。

修复代码

补全缺失的分支逻辑即可,你可以根据自身业务规则调整else分支的返回规则,示例如下:

def get_start_times(x: pd.Series):
    site_hour = x["site_time"].hour
    shift_hour = x["shift_start_time"].hour
    if 0 <= site_hour <= 6:
        return x['shift_start_time'] - pd.Timedelta(days=1)
    elif shift_hour <= site_hour <= 23:
        return x["shift_start_time"]
    # 补全剩余场景的处理逻辑,示例为跨天场景往前推1天,可按需修改
    else:
        return x['shift_start_time'] - pd.Timedelta(days=1)

# apply调用可以直接传函数对象,不需要套一层lambda
df["shift_start_time"] = df.apply(get_start_times, axis=1)
性能优化建议

行级apply是逐行循环实现,数据量较大时执行效率很低,推荐使用pandas原生向量化操作实现相同逻辑,性能可提升数十倍:

# 提取小时字段作为临时判断列
df["site_hour"] = df["site_time"].dt.hour
df["shift_hour"] = df["shift_start_time"].dt.hour

# 定义判断条件
cond_night = df["site_hour"].between(0, 6)
cond_same_day = (df["shift_hour"] <= df["site_hour"]) & (df["site_hour"] <= 23)

# 按条件批量赋值
df.loc[cond_night | ~cond_same_day, "shift_start_time"] -= pd.Timedelta(days=1)

# 删除临时列
df.drop(columns=["site_hour", "shift_hour"], inplace=True)

内容的提问来源于stack exchange,提问作者PandasM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.31 06:15:44