如何用R批量处理德国气象站文本文件:筛选指定时段数据
Got it, let's walk through how to batch process those 1080 German weather station TXT files in R to keep only the 1980-2014 records. This approach is efficient and scalable for large file sets, so you won't have to manually edit each one.
First, get your environment set up so R can find your files and save the filtered versions safely:
# Set working directory to where your raw TXT files are stored (update the path!) setwd("path/to/your/raw_weather_files") # Create a new folder to save filtered data (avoids overwriting original files) dir.create("filtered_1980_2014", showWarnings = FALSE) # Get a list of all TXT files in your folder file_list <- list.files(pattern = "\\.txt$", full.names = TRUE)
Next, we'll build a reusable function that handles reading a single file, cleaning up dates, and filtering the time range. You'll need to tweak this a bit to match your actual data structure—here's a flexible starting point:
filter_weather_data <- function(file_path) { # Read the TXT file (adjust sep/header based on your file's format) # Example: if your file uses commas instead of tabs, change sep to "," weather_data <- read.delim(file_path, sep = "\t", header = TRUE, stringsAsFactors = FALSE) # Convert your date column to a proper Date type (update "date_column" to your actual column name) # Case 1: If date is stored as numeric YYYYMMDD (e.g., 19800101) weather_data$clean_date <- as.Date(as.character(weather_data$date_column), format = "%Y%m%d") # Case 2: If date is already a string like "1980-01-01", use this instead: # weather_data$clean_date <- as.Date(weather_data$date_column) # Filter for the 1980-2014 time frame filtered_data <- subset(weather_data, clean_date >= as.Date("1980-01-01") & clean_date <= as.Date("2014-12-31")) # Define the output file path (keeps original filename in the new folder) output_file <- file.path("filtered_1980_2014", basename(file_path)) # Save the filtered data back to TXT (match separator to your input) write.table(filtered_data, output_file, sep = "\t", row.names = FALSE, quote = FALSE) # Optional: Print progress so you know how things are going cat("Completed:", basename(file_path), "\n") }
Now run the function across all 1080 files. We'll add error handling so a single problematic file doesn't crash the whole process:
# Base R approach with error catching lapply(file_list, function(file) { tryCatch({ filter_weather_data(file) }, error = function(e) { cat("Error processing", basename(file), ":", e$message, "\n") }) }) # Tidyverse alternative (if you prefer purrr) # library(purrr) # walk(file_list, ~tryCatch(filter_weather_data(.x), error = function(e) cat("Issue with", basename(.x), ":", e$message, "\n")))
- File Separator: If your TXT files use commas, spaces, or another delimiter, update the
separgument inread.delimandwrite.table. - Date Format: Double-check how dates are stored in your dataset—use the correct
formatcode inas.Date()(e.g.,%d/%m/%Yfor day/month/year). - Column Names: Replace
"date_column"with the actual name of the date column in your files.
This workflow will automatically process every file, save the filtered versions to a dedicated folder, and leave your original data untouched.
内容的提问来源于stack exchange,提问作者Jerry07

