You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用tz_localize时出现1970-01-01日期异常的技术求助

解决分组后时间索引变为1970年日期的问题

问题重现

用户的分钟级行情DataFrame已将timestamp转为datetime类型,按symbol和日期分组后,执行以下代码时,输出的索引全为1970年代日期:

df['timestamp'] = pd.to_datetime(df['timestamp'])
seperate_days = df.groupby(['symbol', df['timestamp'].dt.date])

for name, ohlc in seperate_days:
    ohlc.set_index('timestamp')
    ohlc.index = pd.to_datetime(ohlc.index)
    ohlc.index = ohlc.index.tz_localize(tz='America/New_York')
    print(ohlc.index)

错误原因

  1. ohlc.set_index('timestamp')默认返回新的DataFrame,未修改原ohlc对象,因此ohlc仍保留原来的整数索引(0、1、2...)。
  2. 将整数索引转为datetime时,这些整数会被解析为从1970-01-01开始的纳秒时间戳,最终生成1970年的日期。
  3. 额外问题:原始timestamp已带UTC时区(+00:00),使用tz_localize会给已有时区的时间强行添加新时区,逻辑错误,应使用tz_convert转换时区。

修正方案

方案1:正确设置索引并转换时区

修改set_index的使用方式,确保索引被正确替换,同时用tz_convert转换时区:

df['timestamp'] = pd.to_datetime(df['timestamp'])
seperate_days = df.groupby(['symbol', df['timestamp'].dt.date])

for name, ohlc in seperate_days:
    # 方法1:重新赋值ohlc,获取设置索引后的新DataFrame
    ohlc = ohlc.set_index('timestamp')
    # 或者方法2:用inplace=True直接修改原对象
    # ohlc.set_index('timestamp', inplace=True)
    
    # 将UTC时区转换为纽约时区
    ohlc.index = ohlc.index.tz_convert('America/New_York')
    print(ohlc.index)

方案2:提前统一处理时区再分组

先将整个DataFrame的时间转换为目标时区,再按symbol和日期分组,循环时直接使用索引:

# 先把UTC时间转换为纽约时区
df['timestamp'] = pd.to_datetime(df['timestamp']).dt.tz_convert('America/New_York')
# 按symbol和纽约时区的日期分组
seperate_days = df.groupby(['symbol', df['timestamp'].dt.date])

for name, ohlc in seperate_days:
    ohlc = ohlc.set_index('timestamp')
    print(ohlc.index)

关键说明

  • tz_localize:给无时区信息的datetime对象添加时区,适用于原始时间不带时区的场景。
  • tz_convert:将已有时区的datetime对象转换到指定时区,适用于原始时间带时区的场景(如本例的UTC时间)。

内容的提问来源于stack exchange,提问作者a7dc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 07:25:17