如何修改R代码实现按城市目录迭代处理海量压缩数据文件
Modified R Code for City-wise Large Data Processing
Below is the adjusted code that processes data one city directory at a time, avoiding loading all files into memory simultaneously:
# Load required libraries once library(data.table) library(dplyr) library(lubridate) # Set base directory base_dir <- "C:/Users/Alexia/Desktop/Data/Test_Gz" setwd(base_dir) # Decompression function (unchanged from original) decompress <- function(file, dest = sub("\\.gz$", "", file)) { src <- gzfile(file, "rb") on.exit(close(src), add = TRUE) dst <- file(dest, "wb") on.exit(close(dst), add = TRUE) BATCH_SIZE <- 10 * 1024^2 repeat { bytes <- readBin(src, raw(), BATCH_SIZE) if (length(bytes) != 0) { writeBin(bytes, dst) } else { break } } invisible(dest) } # Get list of city directories (all subdirectories in base directory) city_dirs <- list.dirs(path = base_dir, full.names = TRUE, recursive = FALSE) # Process each city directory individually for (city_dir in city_dirs) { city_name <- basename(city_dir) cat("Processing city:", city_name, "\n") # 1. Decompress all .gz files in current city directory gz_files <- list.files(path = city_dir, pattern = "\\.gz$", full.names = TRUE) for (gz_file in gz_files) { decompress(gz_file) } # 2. Read and tidy text files for the city txt_files <- list.files(path = city_dir, pattern = "\\.txt$", full.names = TRUE) # Read each file, skip comment block dt_list <- lapply(txt_files, function(file) { lines <- readLines(file) comment_end <- match("*/", lines) fread(file, skip = comment_end) }) # Clean up column names dt_tidied <- lapply(dt_list, function(x) { # Replace ? with degree symbol in temperature column setnames(x, old = "T2 [?C]", new = "T2 [°C]", skip_absent = TRUE) # Replace brackets with parentheses colnames(x) <- gsub("\\[", "(", colnames(x)) colnames(x) <- gsub("\\]", ")", colnames(x)) x }) # Merge all files for the city merged_city <- rbindlist(dt_tidied, fill = TRUE, use.names = TRUE) # 3. Calculate 15-row averages mn <- merged_city %>% separate(`Date/Time`, into = c("Date", "Time"), sep = "T") mnf <- mn %>% as_tibble() %>% group_by(group = as.integer(gl(n(), 15, n()))) %>% summarise(across(everything(), ~ if(mean(is.na(.x)) > 0.5) NA else mean(.x, na.rm = TRUE))) # Save results to city directory output_path <- file.path(city_dir, paste0("Mean_15_", city_name, ".csv")) write.csv(mnf, output_path, row.names = FALSE) cat("Saved results to:", output_path, "\n") # Optional: Delete decompressed .txt files to free disk space file.remove(txt_files) cat("Cleaned up decompressed files for", city_name, "\n") # Clear memory to avoid accumulation rm(dt_list, dt_tidied, merged_city, mn, mnf) gc() } cat("All cities processed successfully!\n")
Key Changes:
- City-wise iteration: Processes one city's directory at a time instead of all files globally
- Targeted file handling: Only interacts with files in the current city directory
- Memory management: Explicitly removes objects from memory and runs garbage collection after each city
- Per-city output: Saves results as separate CSV files named with the city name (e.g.,
Mean_15_DET.csv) - Optional cleanup: Removes decompressed .txt files after processing to save disk space (comment this line if you need to keep the files)
Content of the question originates from Stack Exchange, question author Alexia k Boston
相关产品推荐
相关产品推荐

