You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Pandas pd.cut重复标签报错,实现DataFrame昼夜时段划分

问题描述

有如下简化后的DataFrame:

DateTime      Value        Date   Time   
0  2022-09-18 06:00:00       5.4  18/09/2022  06:00  
1  2022-09-18 07:00:00       6.0  18/09/2022  07:00  
2  2022-09-18 08:00:00       6.5  18/09/2022  08:00  
3  2022-09-18 09:00:00       6.7  18/09/2022  09:00  
8  2022-09-18 14:00:00       7.9  18/09/2022  14:00  
9  2022-09-18 15:00:00       7.8  18/09/2022  15:00  
10 2022-09-18 16:00:00       7.6  18/09/2022  16:00  
11 2022-09-18 17:00:00       6.8  18/09/2022  17:00  
12 2022-09-18 18:00:00       6.4  18/09/2022  18:00   
13 2022-09-18 19:00:00       5.7  18/09/2022  19:00   
14 2022-09-18 20:00:00       4.8  18/09/2022  20:00   
15 2022-09-18 21:00:00       5.4  18/09/2022  21:00   
16 2022-09-18 22:00:00       4.7  18/09/2022  22:00   
17 2022-09-18 23:00:00       4.3  18/09/2022  23:00   
18 2022-09-19 00:00:00       4.1  19/09/2022  00:00    
19 2022-09-19 01:00:00       4.4  19/09/2022  01:00    
22 2022-09-19 04:00:00       3.5  19/09/2022  04:00    
23 2022-09-19 05:00:00       2.8  19/09/2022  05:00    
24 2022-09-19 06:00:00       3.8  19/09/2022  06:00  

需要新增一列period,按以下规则划分时段:

  • 00:00 - 05:00 → night
  • 06:00 - 18:00 → day
  • 19:00 - 23:00 → night

使用pd.cut时因重复标签报错,代码如下:

df['period'] =  pd.cut(pd.to_datetime(df.DateTime).dt.hour,
       bins=[0, 5, 17, 23],
       labels=['night', 'morning', 'night'],
       include_lowest=True)

报错信息:

ValueError: labels must be unique if ordered=True; pass ordered=False for duplicate labels

解决方案

方法1:添加ordered=False参数并修正bins

按照报错提示设置ordered=False允许重复标签,同时修正bins范围以匹配需求:

df['period'] = pd.cut(
    pd.to_datetime(df['DateTime']).dt.hour,
    bins=[-1, 5, 18, 23],  # 用-1确保0点被包含到第一个区间
    labels=['night', 'day', 'night'],
    include_lowest=True,
    ordered=False
)

方法2:用np.select实现直观条件判断

直接通过条件列表和结果列表赋值,避免pd.cut的标签限制:

import numpy as np

hours = pd.to_datetime(df['DateTime']).dt.hour
conditions = [
    hours.between(0, 5),
    hours.between(6, 18),
    hours.between(19, 23)
]
choices = ['night', 'day', 'night']

df['period'] = np.select(conditions, choices, default=np.nan)

方法3:自定义函数结合apply

通过自定义函数明确时段规则,再批量生成列:

import numpy as np

def get_period(hour):
    if 0 <= hour <= 5:
        return 'night'
    elif 6 <= hour <= 18:
        return 'day'
    elif 19 <= hour <= 23:
        return 'night'
    else:
        return np.nan

df['period'] = pd.to_datetime(df['DateTime']).dt.hour.apply(get_period)

内容的提问来源于stack exchange,提问作者Lana.s

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 01:50:25