You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效识别DataFrame日期列中各日期所属的时间段?

高效匹配Datetime到对应时间段的方法

你的核心问题在于原方法使用了双重循环(apply逐行遍历+列表推导遍历所有period),时间复杂度为O(n*m),数据量越大速度越慢。下面提供两种矢量化的高效解决方案,完全避免显式循环:

方案1:针对固定频率的连续时间段(如示例中的日度周期)

如果你的时间段是固定频率的连续区间(比如按天、按小时划分),直接使用pandas的dt.to_period()方法即可,这是最快速的矢量化操作:

import pandas as pd
import datetime

today = datetime.datetime.now().date()
dates = pd.date_range(start=today, freq='1H', periods=100)
df = pd.DataFrame({'Dates': dates})

# 直接将Datetime列转换为日度Period
df['Period'] = df['Dates'].dt.to_period('D')

这个方法的时间复杂度为O(n),内部由pandas的C语言底层实现,处理百万级数据也能瞬间完成。

方案2:针对自定义/非连续时间段

如果你的时间段是不规则、非连续的自定义区间,使用pd.cut结合时间戳分箱来实现高效匹配:

import pandas as pd
import datetime

today = datetime.datetime.now().date()
dates = pd.date_range(start=today, freq='1H', periods=100)
df = pd.DataFrame({'Dates': dates})

periods = pd.period_range(start=today, freq='d', periods=10)

# 将Period的起止时间转换为时间戳,作为分箱边界
bin_starts = [p.start_time.timestamp() for p in periods]
bin_ends = [p.end_time.timestamp() for p in periods]
# 构造完整的分箱边界(左闭右开,包含第一个区间的起始和最后一个区间的结束)
bins = [bin_starts[0]] + bin_ends

# 用pd.cut对Datetime的时间戳进行分箱,得到对应Period的索引
df['period_index'] = pd.cut(
    df['Dates'].dt.timestamp(),
    bins=bins,
    labels=False,
    include_lowest=True  # 确保第一个区间包含起始点
)

# 根据索引映射到对应的Period对象
df['Period'] = periods[df['period_index']].values

这个方法的时间复杂度为O(n log m),相比原方法的双重循环,效率提升几个数量级。

原方法低效的原因

原代码中df.Dates.apply(lambda x: [p for p in periods if (p.start_time<=x)&(p.end_time>x)][0]):

  • apply会逐行处理每个Datetime对象,属于显式循环
  • 每个行又要遍历所有Period做条件判断,形成双重循环
    当数据量达到万级以上时,这种方法的性能会急剧下降。

内容的提问来源于stack exchange,提问作者Charlie_M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 11:24:24