R中绘制股票ARIMA预测值及大数据量绘图解决方案
解决分钟级股票数据绘图问题及小时级聚合预测方案
一、分钟级数据绘图的解决方案
37万+条分钟级数据点远超过屏幕可显示的像素数量,直接绘制所有点不仅视觉上无法区分细节,还会触发绘图性能报错。以下两种方案可以解决这个问题:
方案1:对原始数据采样后绘图
通过间隔采样减少绘图点数量,同时保留预测值进行对比(这里以每10分钟取1个点为例,可根据需求调整采样间隔):
# 为原始数据添加顺序索引(假设数据已按时间排序) data$idx <- seq_len(nrow(data)) # 间隔采样原始数据 sampled_data <- data[seq(1, nrow(data), by = 10), ] # 整理预测数据,与采样数据的索引衔接 forecast_df <- data.frame( idx = max(sampled_data$idx) + seq_len(10), close = forecastedValues$mean, lower = forecastedValues$lower, upper = forecastedValues$upper, type = "Forecast" ) sampled_data$type <- "Original" # 合并绘图数据 plot_data <- rbind( sampled_data[, c("idx", "close", "type")], forecast_df[, c("idx", "close", "type")] ) # ggplot2绘图 ggplot(plot_data, aes(x = idx, y = close, color = type)) + geom_line(size = 0.5) + geom_ribbon(data = forecast_df, aes(ymin = lower, ymax = upper, fill = type), alpha = 0.2) + labs(x = "Time Index", y = "Close Price", title = "Sampled Original Data vs Forecast") + theme_minimal()
方案2:只绘制近期数据+预测
如果只需要观察近期趋势,直接截取最后N个原始数据点(比如最后5000条):
# 截取最后5000条原始数据 recent_data <- tail(data, 5000) recent_data$idx <- seq(nrow(data) - 5000 + 1, nrow(data)) # 整理预测数据 forecast_df <- data.frame( idx = nrow(data) + seq_len(10), close = forecastedValues$mean, lower = forecastedValues$lower, upper = forecastedValues$upper, type = "Forecast" ) recent_data$type <- "Original" # 合并绘图数据 plot_data <- rbind( recent_data[, c("idx", "close", "type")], forecast_df[, c("idx", "close", "type")] ) # 绘图 ggplot(plot_data, aes(x = idx, y = close, color = type)) + geom_line(size = 0.5) + geom_ribbon(data = forecast_df, aes(ymin = lower, ymax = upper, fill = type), alpha = 0.2) + labs(x = "Time Index", y = "Close Price", title = "Recent Original Data vs Forecast") + theme_minimal()
二、转换为小时级数据的完整流程
如果需要分析更宏观的趋势,可以将分钟级数据聚合为小时级,再进行建模和绘图:
1. 处理时间戳并按小时聚合
首先确保你的CSV包含时间戳列(比如timestamp),先解析时间并按小时分组计算平均收盘价:
library(dplyr) library(lubridate) # 读取数据并解析时间戳(根据实际格式调整,比如dmy_hms) data <- read.csv("C:/Users/HP Pavilion/Desktop/Time Series Major/BHARTIARTL__EQ__NSE__NSE__MINUTE.csv") data$timestamp <- ymd_hms(data$timestamp) # 按小时分组,计算每小时平均收盘价(也可用last(close)取小时末收盘价) hourly_data <- data %>% mutate(hour = floor_date(timestamp, "hour")) %>% group_by(hour) %>% summarise(close_mean = mean(close, na.rm = TRUE)) %>% ungroup() # 转换为时间序列对象(每日24小时频率) hourly_ts <- ts(hourly_data$close_mean, frequency = 24)
2. 小时级ARIMA建模与预测
# 自动选择最优ARIMA模型 hourly_arima <- auto.arima(hourly_ts) summary(hourly_arima) # 预测未来24小时的收盘价(可调整h值改变预测时长) hourly_forecast <- forecast(hourly_arima, level = c(95), h = 24)
3. 小时级数据绘图
方式1:使用forecast包自带绘图函数(快捷)
plot(hourly_forecast, main = "Hourly Close Price Forecast", xlab = "Hour", ylab = "Average Close Price")
方式2:使用ggplot2自定义绘图
# 整理原始小时级数据 hourly_plot_data <- data.frame( hour = hourly_data$hour, close = hourly_data$close_mean, type = "Original" ) # 整理预测数据,生成对应的时间戳 forecast_hours <- seq(max(hourly_data$hour), by = "hour", length.out = 25)[-1] hourly_forecast_df <- data.frame( hour = forecast_hours, close = hourly_forecast$mean, lower = hourly_forecast$lower, upper = hourly_forecast$upper, type = "Forecast" ) # 合并数据并绘图 combined_plot_data <- rbind(hourly_plot_data, hourly_forecast_df) ggplot(combined_plot_data, aes(x = hour, y = close, color = type)) + geom_line(size = 0.8) + geom_ribbon(data = hourly_forecast_df, aes(ymin = lower, ymax = upper, fill = type), alpha = 0.2) + labs(x = "Time", y = "Average Close Price", title = "Hourly Original Data vs Forecast") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
内容的提问来源于stack exchange,提问作者Karandeep Saluja
相关产品推荐
相关产品推荐

