You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改R代码实现按城市目录迭代处理海量压缩数据文件

Modified R Code for City-wise Large Data Processing

Below is the adjusted code that processes data one city directory at a time, avoiding loading all files into memory simultaneously:

# Load required libraries once
library(data.table)
library(dplyr)
library(lubridate)

# Set base directory
base_dir <- "C:/Users/Alexia/Desktop/Data/Test_Gz"
setwd(base_dir)

# Decompression function (unchanged from original)
decompress <- function(file, dest = sub("\\.gz$", "", file)) {
  src <- gzfile(file, "rb")
  on.exit(close(src), add = TRUE)
  dst <- file(dest, "wb")
  on.exit(close(dst), add = TRUE)

  BATCH_SIZE <- 10 * 1024^2
  repeat {
    bytes <- readBin(src, raw(), BATCH_SIZE)
    if (length(bytes) != 0) {
      writeBin(bytes, dst)
    } else {
      break
    }
  }

  invisible(dest)
}

# Get list of city directories (all subdirectories in base directory)
city_dirs <- list.dirs(path = base_dir, full.names = TRUE, recursive = FALSE)

# Process each city directory individually
for (city_dir in city_dirs) {
  city_name <- basename(city_dir)
  cat("Processing city:", city_name, "\n")

  # 1. Decompress all .gz files in current city directory
  gz_files <- list.files(path = city_dir, pattern = "\\.gz$", full.names = TRUE)
  for (gz_file in gz_files) {
    decompress(gz_file)
  }

  # 2. Read and tidy text files for the city
  txt_files <- list.files(path = city_dir, pattern = "\\.txt$", full.names = TRUE)
  
  # Read each file, skip comment block
  dt_list <- lapply(txt_files, function(file) {
    lines <- readLines(file)
    comment_end <- match("*/", lines)
    fread(file, skip = comment_end)
  })

  # Clean up column names
  dt_tidied <- lapply(dt_list, function(x) {
    # Replace ? with degree symbol in temperature column
    setnames(x, old = "T2 [?C]", new = "T2 [°C]", skip_absent = TRUE)
    # Replace brackets with parentheses
    colnames(x) <- gsub("\\[", "(", colnames(x))
    colnames(x) <- gsub("\\]", ")", colnames(x))
    x
  })

  # Merge all files for the city
  merged_city <- rbindlist(dt_tidied, fill = TRUE, use.names = TRUE)

  # 3. Calculate 15-row averages
  mn <- merged_city %>% separate(`Date/Time`, into = c("Date", "Time"), sep = "T")
  mnf <- mn %>% 
    as_tibble() %>%
    group_by(group = as.integer(gl(n(), 15, n()))) %>%
    summarise(across(everything(), ~ if(mean(is.na(.x)) > 0.5) NA else mean(.x, na.rm = TRUE)))

  # Save results to city directory
  output_path <- file.path(city_dir, paste0("Mean_15_", city_name, ".csv"))
  write.csv(mnf, output_path, row.names = FALSE)
  cat("Saved results to:", output_path, "\n")

  # Optional: Delete decompressed .txt files to free disk space
  file.remove(txt_files)
  cat("Cleaned up decompressed files for", city_name, "\n")

  # Clear memory to avoid accumulation
  rm(dt_list, dt_tidied, merged_city, mn, mnf)
  gc()
}

cat("All cities processed successfully!\n")

Key Changes:

  • City-wise iteration: Processes one city's directory at a time instead of all files globally
  • Targeted file handling: Only interacts with files in the current city directory
  • Memory management: Explicitly removes objects from memory and runs garbage collection after each city
  • Per-city output: Saves results as separate CSV files named with the city name (e.g., Mean_15_DET.csv)
  • Optional cleanup: Removes decompressed .txt files after processing to save disk space (comment this line if you need to keep the files)

Content of the question originates from Stack Exchange, question author Alexia k Boston

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 20:14:55