You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

合并30分钟间隔DataFrame时午夜数据丢失问题排查与解决

解决多CSV合并时午夜时间戳数据丢失/重复问题

问题场景

  • 多份CSV数据,每30分钟生成1条含23个变量的记录,单文件覆盖24小时
  • 已完成第一步:基于TIMESTAMP列合并所有CSV得到dataMerge,此步骤无异常
  • 需完成第二步:创建包含2024全年30分钟时间戳的数据集df(共17521条观测),与dataMerge合并以生成含缺失时段的完整全年数据集
  • 合并时遇到异常:
    • 常规合并后所有午夜行数据丢失
    • 使用all=T参数时,午夜行出现重复(一行有数据、一行无数据)

问题原因

CSV文件中的TIMESTAMP列以字符串格式存储,未转换为标准的POSIXct时间类型,导致与生成的POSIXct格式时间戳匹配时出现异常,尤其是午夜时间的格式兼容问题。

修正后的完整代码

FileNames <- list.files(path = path_load)
headers <- read.csv(file = paste(path_load, FileNames[1], sep=""), skip = 1, header = F, nrows = 1, as.is = T)

FirstFile <- read.csv(file = paste(path_load, FileNames[1], sep=""), skip = 4, header = F)
colnames(FirstFile) = headers
FirstFile$TIMESTAMP <- as.POSIXct(FirstFile$TIMESTAMP,format="%Y-%m-%d %H:%M:%S")

SecondFile <- read.csv(file = paste(path_load, FileNames[2], sep=""), skip = 4, header = F)
colnames(SecondFile) = headers
SecondFile$TIMESTAMP <- as.POSIXct(SecondFile$TIMESTAMP,format="%Y-%m-%d %H:%M:%S")

dataMerge <- merge(FirstFile, SecondFile, all = T)

for(i in 3:length(FileNames)){
  ReadInMerge <- read.csv(file=paste(path_load, FileNames[i], sep=""), skip = 4, header = F)
  colnames(ReadInMerge) = headers
  ReadInMerge$TIMESTAMP <- as.POSIXct(ReadInMerge$TIMESTAMP,format="%Y-%m-%d %H:%M:%S")
  dataMerge <- merge(dataMerge, ReadInMerge, all=T)
}

df <- data.frame(TIMESTAMP = seq(as.POSIXct("2024-01-01", tz = "EST"),
                                as.POSIXct("2024-12-31", tz = "EST"),
                                by=(30*60)))

AllMerge <- merge(df, dataMerge, all.x = T, by = "TIMESTAMP")

内容的提问来源于stack exchange,提问作者LPs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:22:58