R Studio批量处理50个大型TXT文件的自动化代码求助
Got it, let's turn your single-file workflow into an efficient batch system that handles all 50 of your large TXT files and combines their results into one target CSV. Here's a streamlined solution that keeps your existing logic intact while optimizing for speed (critical with 100MB files):
Step 1: Set Up Environment & Paths
First, load your required packages and define clear paths for your input files and output target. This keeps your code easy to adjust later.
library(data.table) library(dplyr) library(tidyverse) # Folder where all your TXT files are stored input_folder <- "C:/Users/shegu/Desktop/" # Path for your final combined output target_path <- "C:/Users/shegu/Desktop/BSE_Data/Target.csv"
Step 2: Get List of All TXT Files
Use list.files() to grab only the .txt files in your input folder—this avoids accidentally processing other file types.
# Fetch full paths to all TXT files txt_files <- list.files(path = input_folder, pattern = "\\.txt$", full.names = TRUE)
Step 3: Create a Reusable Processing Function
Wrap your single-file logic into a function. This lets you apply the exact same steps to every file, and keeps your code modular and easy to debug. The function will return the grouped summary table so we can combine results later.
process_single_file <- function(file_path) { # Read the file (fread is way faster for large datasets) myfile <- fread(file_path, sep = "|", header = FALSE, stringsAsFactors = TRUE) # Rename columns to match your existing setup colnames(myfile) <- c( "Trading_Session", "Scrip_Code", "Buy_Sell", "Order_Type", "Rate_in_Paise", "Quantity", "Avl_Quantity", "Order_Time_Stamp", "Retention", "AUD_Code", "Order_ID", "Action_ID", "Error_Code", "ALGO_Flag" ) # Format columns efficiently using data.table's in-place operator myfile[, `:=`( Order_Time_Stamp = as.Date(Order_Time_Stamp, "%Y-%m-%d %H:%M:%S"), Scrip_Code = as.factor(Scrip_Code), Order_ID = as.factor(Order_ID) )] # Perform group-by summary (added a clear column name for the count) myfile_by_AUD_Code <- myfile %>% group_by(Scrip_Code, ALGO_Flag, AUD_Code) %>% summarise(record_count = n(), .groups = "drop") return(myfile_by_AUD_Code) }
Note: I used data.table's := operator for column formatting—it modifies data in place, which saves memory and time with large files. I also renamed the n() output to record_count for readability.
Step 4: Batch Process & Combine Results
Use map_dfr() from purrr (part of tidyverse) to apply the function to every file and automatically bind all results into a single data frame. For even faster performance with massive datasets, you can swap in data.table's rbindlist().
# Process all files and merge results into one table all_results <- map_dfr(txt_files, process_single_file) # Alternative (even faster for very large data): # all_results <- rbindlist(lapply(txt_files, process_single_file))
Step 5: Write Combined Results to CSV
Finally, export the merged summary data to your target CSV file. This will contain the grouped stats from all 50 input files in one place.
write.csv(all_results, target_path, row.names = FALSE)
Quick Tips for Large Files:
- Stick with
fread—it’s significantly faster than base R’sread.csvfor big datasets. - If you hit memory limits, you can process files in smaller batches or write intermediate results incrementally, but this setup should work for most modern systems with 50x100MB files.
内容的提问来源于stack exchange,提问作者Abhishek G

