R语言新手求助:文件处理函数内存占用过高与性能优化
Fixing Slow Performance & High Memory Usage in Your R Function
Hey there! Let's break down why your function is taking so long to handle a small 2.7MB zip file—there are a few key inefficiencies in the original code that we can fix easily.
Key Issues in the Original Code
- Inefficient file movement: Using
file.copy()followed byfile.remove()creates duplicate files temporarily, which wastes memory and time. A direct rename is much faster. - Redundant loop for deleting directories: Your for loop tries to delete all subdirectories every iteration, which is unnecessary and slows things down.
- Unoptimized file listing: The way you're checking for files in subfolders can be streamlined to avoid extra processing.
Optimized Version of Your Function
create_dir <- function(directory) { # Create target directory if it doesn't exist (add recursive for robustness) path <- file.path(getwd(), directory) if (!file.exists(path)) { dir.create(path, recursive = TRUE) } # Download and unzip the file link <- "https://d396qusza40orc.cloudfront.net/rprog%2Fdata%2Fspecdata.zip" temp <- tempfile() # Specify method for faster download (use "wininet" on Windows if curl isn't available) download.file(link, temp, mode = "wb", method = "curl") unzip(temp, exdir = path) unlink(temp) # Get immediate subdirectories under the target path sub_dirs <- list.dirs(path, recursive = FALSE) # Move files from subdirectories to main directory and clean up for (sub_dir in sub_dirs) { # List all files in the subfolder with full paths sub_files <- list.files(sub_dir, full.names = TRUE) # Move files directly (no copy+remove = way faster, less memory) file.rename(sub_files, file.path(path, basename(sub_files))) # Delete the now-empty subfolder right away unlink(sub_dir, recursive = TRUE) } } # Test the optimized function create_dir("testDirectory")
What Changed & Why
- Robust directory creation: Added
recursive = TRUEtodir.create()so it creates missing parent directories automatically, no extra checks needed. - Faster downloads: Specifying
method = "curl"(orwininetfor Windows) often speeds up file transfers compared to the default method in some environments. - Direct file movement:
file.rename()moves files without duplicating them, which cuts memory usage and processing time drastically—no more copying then deleting files. - Streamlined subfolder handling: We only target immediate subfolders (instead of all recursive paths), process each one in a single loop, and delete empty folders right after moving files. No redundant iterations here.
- Removed unnecessary loops: The original code looped through every file and then every directory separately; we simplified this to one focused loop over subfolders.
This version should run in seconds instead of minutes, even for small files, and use far less memory. Give it a try!
内容的提问来源于stack exchange,提问作者Stephen Aung
相关产品推荐
相关产品推荐

