求助:使用R语言lapply函数生成符合需求的CSV表格
Hey there! Since you're new to R and cramming to get this time-series data task done in a hurry, let's walk through exactly how to use lapply to pull together your hourly data into a single CSV. I'll keep this step-by-step and beginner-friendly.
Step 1: Prep Your Environment & File Paths
First, let's make sure you have the right tools. The tidyverse package will make reading and cleaning data way easier—if you haven't installed it yet, run the install line first. Then we'll grab all the file paths for your monthly data files.
# Install tidyverse if you haven't already (remove the # to run) # install.packages("tidyverse") library(tidyverse) # Replace this with the actual folder where your data files live data_folder <- "C:/your_data_folder_here" # Get all file paths matching your naming pattern # Adjust the `pattern` to match how your files are named (e.g., "temp_2000_01.csv") # Example pattern for files like "location_year_month.csv": ".*_\\d{4}_\\d{2}\\.csv" file_paths <- list.files( path = data_folder, pattern = ".*_\\d{4}_\\d{2}\\.csv", # Tweak this to match your file names! full.names = TRUE # Keeps the full file path instead of just the filename )
Step 2: Write a Function to Read Single Files
We need a reusable function that reads one file, extracts your target variable, and adds context like location, year, and month (since these are tied to the file itself). Adjust this based on your actual file structure:
read_single_month <- function(file_path) { # Extract location, year, month from the filename (tweak this to match your naming!) file_name <- basename(file_path) # Gets just the filename without the folder path name_parts <- str_split(file_name, "_", simplify = TRUE) # Split by underscores # Example: if filename is "nyc_2000_01.csv", name_parts[1] = "nyc", [2] = "2000", [3] = "01.csv" location <- name_parts[1] year <- as.integer(name_parts[2]) month <- as.integer(str_remove(name_parts[3], "\\.csv")) # Remove .csv extension # Read the file and keep only your target variable (replace "your_target_variable"!) # Add a datetime column if you have one (critical for hourly data!) monthly_data <- read_csv(file_path) %>% select(datetime, your_target_variable) # Replace with your actual column names mutate( location = location, year = year, month = month ) return(monthly_data) }
Pro tip: If your files already include a datetime column with year/month/hour, you can skip extracting year/month from the filename—just use lubridate::year(datetime) and lubridate::month(datetime) instead.
Step 3: Use lapply to Batch Process All Files
Now we'll feed all our file paths into lapply, which will run our read_single_month function on every file automatically:
# Read all files into a list of data frames all_data_list <- lapply(file_paths, read_single_month) # Merge the list into one big data frame combined_data <- bind_rows(all_data_list) # Tidyverse way, or use do.call(rbind, all_data_list)
Step 4: Save the Final CSV
Finally, write your combined dataset to a CSV file:
write_csv(combined_data, "combined_hourly_data_20years.csv")
Quick Troubleshooting Tips
- File naming issues: If your filenames don't follow the pattern I used, adjust the
str_splitpart to match how your files are labeled (e.g., if location is at the end of the filename, usename_parts[length(name_parts)-2]). - Reading errors: If some files fail to read, wrap the code inside
read_single_monthin atryCatchto catch errors without breaking the whole process:read_single_month <- function(file_path) { tryCatch({ # All the existing code here }, error = function(e) { message(paste("Oops, failed to read", file_path, ":", e$message)) return(NULL) }) } # Then filter out failed files before merging: all_data_list <- Filter(Negate(is.null), all_data_list) - Non-CSV files: If your data is in NetCDF or another format, swap
read_csvwith the appropriate function (e.g.,ncdf4::ncvar_getfor NetCDF) and adjust the variable extraction logic.
Let me know if you need to tweak any part to fit your exact file setup!
内容的提问来源于stack exchange,提问作者Orlin_Mitov

