生成伪数据分析:大文件脚本无响应及data.table fread报错问题
Hey there! Let's work through your problem with handling that massive dataset (1000 rows, 8.6M columns) and fixing the fread error, plus efficiently shuffling the first 6 columns as you need.
Step 1: Fix the fread Crash Error
The error you're seeing (possible actions: 1: abort (with core dump)) is almost certainly due to either:
- Insufficient system memory to load such a wide dataset (8.6 million columns is extremely large—even 1000 rows of integers would take ~3.4GB of RAM)
- Automatic parsing issues from
freadstruggling with the sheer number of columns
Try these adjusted fread parameters to resolve the crash:
library(data.table) library(parallel) # Read the big file with optimized settings dt <- fread( "bigDataFile.txt", sep = "\t", # Replace with your actual delimiter (e.g., "," for CSV) header = TRUE, # Set to FALSE if your file has no column headers nThread = detectCores(), # Use all available CPU cores for faster reading showProgress = FALSE, # Disable progress bar to reduce overhead integer64 = "character", # Avoid parsing issues with large integers verbose = FALSE # Disable verbose output to speed things up )
If you still hit memory issues, check your system's available RAM—you might need to close other applications or use a machine with more memory. For extreme cases, you could read the file in chunks by columns, but with only 1000 rows, it's better to prioritize getting enough RAM first.
Step 2: Efficiently Shuffle the First 6 Columns & Merge
Once you've successfully loaded the data, use data.table's fast operations to shuffle the first 6 columns and combine them with the rest. Here are two common scenarios based on what you mean by "shuffle the first 6 columns":
Scenario 1: Shuffle the order of the first 6 columns
If you want to randomize the sequence of the first 6 columns (e.g., reorder columns 1-6 randomly) while keeping all rows intact and appending columns 7 onwards:
# Generate a random order for the first 6 columns shuffled_col_indices <- sample(1:6) # Create the shuffled dataset: reordered first 6 columns + remaining columns dt_shuffled <- cbind(dt[, ..shuffled_col_indices], dt[, 7:ncol(dt)])
Scenario 2: Shuffle the rows of the first 6 columns (keep remaining columns' row order)
If you want to randomly reorder the rows of the first 6 columns (so each row in columns 1-6 is from a random original row) while leaving columns 7+ in their original row order:
# Generate a random order of row indices shuffled_row_indices <- sample(1:nrow(dt)) # Combine shuffled first 6 columns with unchanged remaining columns dt_shuffled <- cbind(dt[shuffled_row_indices, 1:6], dt[, 7:ncol(dt)])
Step 3: Save the Result
Finally, write the shuffled dataset back to a file efficiently with fwrite (data.table's fast write function):
fwrite(dt_shuffled, "shuffled_bigDataFile.txt", sep = "\t")
内容的提问来源于stack exchange,提问作者star

