You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas处理多时间框架OHLCV数据?代码NaN问题求助

解决Pandas处理多时间框架OHLCV数据的NaN问题

问题背景

需要用Pandas处理5分钟、15分钟、30分钟、1小时、4小时、1天的OHLCV数据,当前代码能创建正确的列结构,但输出存在大量NaN值,插值数据未正确填充。

问题根源

  1. 索引不匹配:原代码中,resample后的结果是对应时间框架的起始时间索引,而原DataFrame是分钟级索引,直接赋值时只有极少数索引能匹配,导致大部分位置为NaN。
  2. 时间清理逻辑错误:clean_data中修改索引的方式会导致时间被错误偏移,后续resample的基础数据存在问题。
  3. 空列初始化方式错误:直接创建空Series时未对齐原索引,后续赋值无法正确填充。

修复后的完整代码

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

def load_data(file_path: str) -> pd.DataFrame:
    """加载CSV格式的OHLCV数据"""
    # 指定列名确保数据结构正确,避免文件无表头的问题
    names = ['time','open','close','high','low','volume']
    df = pd.read_csv(file_path, parse_dates=['time'], index_col='time', names=names, skiprows=1)
    return df

def clean_data(df: pd.DataFrame) -> pd.DataFrame:
    """清理数据,处理时间间隙与重复值"""
    # 先排序索引,保证时间顺序正确
    df = df.sort_index()
    
    # 移除重复的时间戳,保留第一条
    df = df.drop_duplicates(keep='first')
    
    # 按分钟重采样并向前填充,补全时间间隙
    # 仅对价格和成交量列做填充,确保数据连续
    df = df.resample('T').ffill()
    
    return df

def resample_data(df: pd.DataFrame) -> pd.DataFrame:
    """生成多时间框架的OHLCV数据,合并到原DataFrame"""
    time_frames = {
        '5_min': '5T',
        '15_min': '15T',
        '30_min': '30T',
        '1_hour': '60T',
        '4_hour': '240T',
        '1_day': '1440T'
    }
    
    columns_to_resample = ['open', 'high', 'low', 'close', 'volume']
    
    # 遍历每个时间框架,生成重采样数据后合并到原DataFrame
    for col_name, freq in time_frames.items():
        resampled = df[columns_to_resample].resample(freq).agg({
            'open': 'first',
            'high': 'max',
            'low': 'min',
            'close': 'last',
            'volume': 'sum'
        })
        # 插值填充重采样后的NaN(比如成交量可能存在的空值)
        resampled = resampled.interpolate(method='linear')
        
        # 给重采样后的列添加前缀,避免和原列冲突
        resampled.columns = [f"{col_name}_{subcol}" for subcol in resampled.columns]
        
        # 将重采样数据合并到原DataFrame,使用向前填充补全分钟级索引的缺失值
        df = df.merge(resampled, left_index=True, right_index=True, how='left')
        df = df.ffill()
    
    # 检查NaN值情况
    print("NaN值统计:")
    print(df.isna().sum())
    
    return df

# 加载数据
file_path = r'C:\Users\Shadow\.cursor-tutor\projects\Machine Learning Modules\btcusd_ISO8601.csv'
df = load_data(file_path)

# 清理数据
df = clean_data(df)

# 生成多时间框架数据
df = resample_data(df)

# 查看数据前几行
print("\n数据预览:")
print(df.head())

# 绘制收盘价曲线
plt.figure(figsize=(15, 5))
plt.plot(df['close'], label='原始分钟收盘价')
plt.plot(df['1_day_close'], label='日线收盘价', linewidth=2)
plt.title('比特币价格走势', fontsize=15)
plt.ylabel('价格(美元)')
plt.legend()
plt.savefig('bitcoin_price_trend.png')
plt.show()

关键修改说明

  • 数据加载优化:指定列名并跳过表头(如果文件有表头),确保数据列匹配正确。
  • 清理逻辑修正:调整时间处理顺序,先排序去重再重采样填充,避免时间索引被错误偏移。
  • 重采样合并方式改进:
    • 使用字典映射时间框架和列名,代码更清晰;
    • 给重采样后的列添加前缀(如5_min_open),替代原多级列结构,更易操作;
    • 通过merge合并重采样数据,再用ffill将时间框架的数值填充到所有分钟级索引,解决NaN问题;
  • 插值逻辑保留:对重采样后的少量NaN值(如成交量)仍用线性插值填充。

内容的提问来源于stack exchange,提问作者Omar Bakri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 10:32:34