You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于文本标签按半小时间隔分割大型CSV文件

拆分带分隔标签的CSV文件为半小时数据块

R语言实现方法

  1. 读取文件所有行,避免分隔标签干扰数据解析
  2. 定位分隔行位置,拆分出48个半小时数据块
  3. 为每个块生成合法文件名并保存,优先使用时段首行的Time值
# 读取目标CSV的所有行
all_lines <- readLines("your_large_file.csv")

# 提取表头(默认第一行为表头)
header <- all_lines[1]

# 定位包含分隔标签的行索引
separator_indices <- grep("zzzz|system sleep", all_lines)

# 计算每个数据块的起止索引
block_starts <- c(2, separator_indices + 1)
block_ends <- c(separator_indices - 1, length(all_lines))

# 截取单日对应的48个数据块
block_starts <- block_starts[1:48]
block_ends <- block_ends[1:48]

# 创建输出文件夹(已存在则不报错)
dir.create("half_hour_blocks", showWarnings = FALSE)

# 循环处理每个数据块
for (i in 1:48) {
  # 拼接当前块的表头和数据行
  block_lines <- c(header, all_lines[block_starts[i]:block_ends[i]])
  
  # 转换为数据框以便提取Time值
  block_df <- read.csv(text = block_lines)
  
  # 生成文件名:优先用Time列首值,否则用序号
  if ("Time" %in% colnames(block_df) && nrow(block_df) > 0) {
    file_name <- paste0(block_df$Time[1], ".csv")
  } else {
    file_name <- paste0("block_", sprintf("%02d", i), ".csv")
  }
  
  # 清理文件名中的非法字符(如冒号、斜杠)
  file_name <- gsub("[:/\\\\]", "_", file_name)
  
  # 保存CSV文件
  write.csv(block_df, file.path("half_hour_blocks", file_name), row.names = FALSE)
}

Python语言实现方法

  1. 读取文件所有行,识别分隔标签行
  2. 拆分出48个数据块,每个块保留表头
  3. 生成合规文件名并保存,优先使用时段首行的Time值
import os
import pandas as pd

# 读取目标CSV的所有行
with open("your_large_file.csv", "r", encoding="utf-8") as f:
    all_lines = f.readlines()

# 提取表头行
header = all_lines[0].strip()

# 定位包含分隔标签的行索引
separator_indices = []
for idx, line in enumerate(all_lines):
    if "zzzz" in line or "system sleep" in line:
        separator_indices.append(idx)

# 计算每个数据块的起止索引
block_starts = [1] + [idx + 1 for idx in separator_indices]
block_ends = [idx - 1 for idx in separator_indices] + [len(all_lines) - 1]

# 截取单日对应的48个数据块
block_starts = block_starts[:48]
block_ends = block_ends[:48]

# 创建输出文件夹(已存在则不报错)
os.makedirs("half_hour_blocks", exist_ok=True)

# 循环处理每个数据块
for i in range(48):
    start_idx = block_starts[i]
    end_idx = block_ends[i]
    
    # 拼接当前块的表头和数据行
    block_content = [header] + [line.strip() for line in all_lines[start_idx:end_idx+1]]
    # 转换为DataFrame
    block_df = pd.read_csv(pd.io.common.StringIO("\n".join(block_content)))
    
    # 生成文件名:优先用Time列首值,否则用序号
    if "Time" in block_df.columns and not block_df.empty:
        file_name = f"{block_df['Time'].iloc[0]}.csv"
    else:
        file_name = f"block_{i+1:02d}.csv"
    
    # 清理文件名中的非法字符
    file_name = file_name.replace(":", "_").replace("/", "_").replace("\\", "_")
    
    # 保存CSV文件
    block_df.to_csv(os.path.join("half_hour_blocks", file_name), index=False)

内容的提问来源于stack exchange,提问作者Nacho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 17:46:27