如何通过循环同时读取多文件夹CSV文件并分配至不同对象
Hey there! Let's get your code sorted for reading those CSVs across four folders, and also flag a big efficiency issue you might not have caught yet. First, let's fix a few small syntax mistakes in your original code:
- R uses
list.files()to get file listings, notlist.dir() - For Windows paths, use forward slashes
/or double backslashes\\to avoid escape character headaches - R is case-sensitive, so it's
stringsAsFactors, notstringasfactors
Corrected & Complete Nested Loop Code
Here's the polished version of your nested loop, with proper file reading and storage (so you don't end up with thousands of loose variables cluttering your environment):
# Define your folder paths (fixed path format for cross-OS compatibility) milk_loc <- "C:/Desktop/milk" cereal_loc <- "C:/Desktop/cereal" OJ_loc <- "C:/Desktop/OJ" cat_loc <- "C:/Desktop/cat" # Grab only CSV files from each folder (filters out non-CSV clutter) milk_files <- list.files(milk_loc, pattern = "\\.csv$", full.names = FALSE) cereal_files <- list.files(cereal_loc, pattern = "\\.csv$", full.names = FALSE) OJ_files <- list.files(OJ_loc, pattern = "\\.csv$", full.names = FALSE) cat_files <- list.files(cat_loc, pattern = "\\.csv$", full.names = FALSE) # Initialize a list to store all file combinations (cleaner than loose variables) all_data <- list() counter <- 1 # Loop through every combination of files for(milk_sheet_name in milk_files){ # Read the current milk CSV (using file.path() for OS-agnostic path building) milk_sheet <- read.csv(file.path(milk_loc, milk_sheet_name), stringsAsFactors = FALSE) for(cereal_file_name in cereal_files){ cereal_file <- read.csv(file.path(cereal_loc, cereal_file_name), stringsAsFactors = FALSE) for(OJ_cup_name in OJ_files){ OJ_cup <- read.csv(file.path(OJ_loc, OJ_cup_name), stringsAsFactors = FALSE) for(cat_paw_name in cat_files){ cat_paw <- read.csv(file.path(cat_loc, cat_paw_name), stringsAsFactors = FALSE) # Store the four data frames as a single entry in our list all_data[[counter]] <- list( milk = milk_sheet, cereal = cereal_file, oj = OJ_cup, cat = cat_paw, combo_label = paste(milk_sheet_name, cereal_file_name, OJ_cup_name, cat_paw_name, sep = "_") ) # If you want to call your target function right here, add it: # your_target_function(milk_sheet, cereal_file, OJ_cup, cat_paw) counter <- counter + 1 } } } }
A Critical Warning:
Wait a second—four folders with 400 files each means your nested loop would run 2.56 BILLION times (400^4). That's impossible to run in any reasonable time, and would eat up all your memory instantly. I'm guessing this isn't actually what you want!
The Far More Practical Alternative
If your files are matched by filename (e.g., milk_001.csv pairs with cereal_001.csv, OJ_001.csv, and cat_001.csv), use a single loop instead. This makes way more sense and is completely manageable:
# Assuming all folders have the same number of matching files file_count <- length(milk_files) matched_data <- list() for(i in 1:file_count){ # Read corresponding files from each folder milk_data <- read.csv(file.path(milk_loc, milk_files[i]), stringsAsFactors = FALSE) cereal_data <- read.csv(file.path(cereal_loc, cereal_files[i]), stringsAsFactors = FALSE) oj_data <- read.csv(file.path(OJ_loc, OJ_files[i]), stringsAsFactors = FALSE) cat_data <- read.csv(file.path(cat_loc, cat_files[i]), stringsAsFactors = FALSE) # Store the matched set in our list matched_data[[i]] <- list( milk = milk_data, cereal = cereal_data, oj = oj_data, cat = cat_data, group_id = gsub("\\.csv$", "", milk_files[i]) # Extract the shared filename prefix ) # Call your target function with the matched data # your_target_function(milk_data, cereal_data, oj_data, cat_data) }
Extra Tips for Better Performance:
- Use
file.path()instead ofpaste0()to build paths—it works on Windows, Mac, and Linux without changes - For large CSVs, swap
read.csv()withdata.table::fread()—it's way faster and uses less memory - If you need speed, use the
foreachpackage withdoParallelto read files in parallel
内容的提问来源于stack exchange,提问作者longlivebrew

