R语言scan()函数替代方案:无需临时文件处理hello开头文本
Got it, let's fix this efficiently without those slow temporary files. The core issue here is we can handle the text manipulation and splitting entirely in memory, cutting out the intermediate file step. Let's break this down:
Step 1: Correctly Target Your Files
First, let's adjust the file listing to get full paths directly (so we don't have to concatenate later) and exclude directories to avoid errors:
maindir <- "/path/to/your/directory" # Replace with your actual directory files <- list.files( path = maindir, pattern = "^hello", # Ensures files start with "hello" full.names = TRUE, recursive = TRUE, include.dirs = FALSE # Skip directories, we only want files )
Step 2: Process Each File In-Memory
For each file, we'll read its content, clean the periods after letters, then split on remaining periods (the ones acting as delimiters) — all without writing to disk. Here's how:
for (file_path in files) { # Read all lines of the file, suppress warnings about incomplete final lines file_content <- readLines(file_path, warn = FALSE) # Combine lines into a single string (handles line breaks seamlessly) combined_content <- paste(file_content, collapse = " ") # Remove periods that follow letters (matches your example: "apple. 10." → "apple 10.") # The regex targets one or more letters followed by a period, replacing with just the letters cleaned_content <- gsub("([A-Za-z]+)\\.", "\\1", combined_content) # Split the cleaned content on periods (matching the original scan(sep=".") behavior) split_result <- strsplit(cleaned_content, "\\.")[[1]] # Optional: Trim whitespace from each split element (common after splitting) split_result <- trimws(split_result) # ------------------------------ # Here you can use split_result for whatever you need next! # Example: Print the result for verification cat("Processing file:", file_path, "\n") print(split_result) }
Key Fixes & Improvements:
- No temporary files: All manipulation happens in memory, which is way faster than writing/reading disk files.
- Accurate regex: The
gsubtargets only periods after letters (not numbers), which matches your requirement to keep periods that act as delimiters (like10.stays intact). - Replaces
scanwithstrsplit: We replicate thescan(sep=".")behavior directly on the cleaned content string, avoiding the need to read from a temporary file. - Cleaner file handling: Using
full.names=TRUEsimplifies file paths, andinclude.dirs=FALSEprevents accidental directory processing.
Why Your Original strsplit Might Have Failed:
If you tried strsplit before without success, it was likely either:
- You ran it on the uncleaned content (with extra periods after letters), leading to unexpected splits.
- Your regex for splitting didn't account for whitespace around periods (the
trimwsstep fixes this by cleaning up extra spaces after splitting).
If you also need to overwrite the original files with the cleaned text (instead of just splitting), add this line after cleaned_content:
writeLines(cleaned_content, file_path)
内容的提问来源于stack exchange,提问作者JMG6V

