You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R Studio批量处理50个大型TXT文件的自动化代码求助

Batch Processing Large TXT Files in RStudio

Got it, let's turn your single-file workflow into an efficient batch system that handles all 50 of your large TXT files and combines their results into one target CSV. Here's a streamlined solution that keeps your existing logic intact while optimizing for speed (critical with 100MB files):

Step 1: Set Up Environment & Paths

First, load your required packages and define clear paths for your input files and output target. This keeps your code easy to adjust later.

library(data.table)
library(dplyr)
library(tidyverse)

# Folder where all your TXT files are stored
input_folder <- "C:/Users/shegu/Desktop/"
# Path for your final combined output
target_path <- "C:/Users/shegu/Desktop/BSE_Data/Target.csv"

Step 2: Get List of All TXT Files

Use list.files() to grab only the .txt files in your input folder—this avoids accidentally processing other file types.

# Fetch full paths to all TXT files
txt_files <- list.files(path = input_folder, pattern = "\\.txt$", full.names = TRUE)

Step 3: Create a Reusable Processing Function

Wrap your single-file logic into a function. This lets you apply the exact same steps to every file, and keeps your code modular and easy to debug. The function will return the grouped summary table so we can combine results later.

process_single_file <- function(file_path) {
  # Read the file (fread is way faster for large datasets)
  myfile <- fread(file_path, sep = "|", header = FALSE, stringsAsFactors = TRUE)
  
  # Rename columns to match your existing setup
  colnames(myfile) <- c(
    "Trading_Session", "Scrip_Code", "Buy_Sell", "Order_Type", 
    "Rate_in_Paise", "Quantity", "Avl_Quantity", "Order_Time_Stamp", 
    "Retention", "AUD_Code", "Order_ID", "Action_ID", "Error_Code", "ALGO_Flag"
  )
  
  # Format columns efficiently using data.table's in-place operator
  myfile[, `:=`(
    Order_Time_Stamp = as.Date(Order_Time_Stamp, "%Y-%m-%d %H:%M:%S"),
    Scrip_Code = as.factor(Scrip_Code),
    Order_ID = as.factor(Order_ID)
  )]
  
  # Perform group-by summary (added a clear column name for the count)
  myfile_by_AUD_Code <- myfile %>% 
    group_by(Scrip_Code, ALGO_Flag, AUD_Code) %>% 
    summarise(record_count = n(), .groups = "drop")
  
  return(myfile_by_AUD_Code)
}

Note: I used data.table's := operator for column formatting—it modifies data in place, which saves memory and time with large files. I also renamed the n() output to record_count for readability.

Step 4: Batch Process & Combine Results

Use map_dfr() from purrr (part of tidyverse) to apply the function to every file and automatically bind all results into a single data frame. For even faster performance with massive datasets, you can swap in data.table's rbindlist().

# Process all files and merge results into one table
all_results <- map_dfr(txt_files, process_single_file)

# Alternative (even faster for very large data):
# all_results <- rbindlist(lapply(txt_files, process_single_file))

Step 5: Write Combined Results to CSV

Finally, export the merged summary data to your target CSV file. This will contain the grouped stats from all 50 input files in one place.

write.csv(all_results, target_path, row.names = FALSE)

Quick Tips for Large Files:

  • Stick with fread—it’s significantly faster than base R’s read.csv for big datasets.
  • If you hit memory limits, you can process files in smaller batches or write intermediate results incrementally, but this setup should work for most modern systems with 50x100MB files.

内容的提问来源于stack exchange,提问作者Abhishek G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 07:52:44