使用R语言readlines从大文本文件随机提取10%条目
Hey there! Let's break down how to solve your problem—whether you're looking to sample 10% of lines (entries) or 10% of words from your large TXT file using R's readLines() function. Here are two scenarios tailored to your needs:
Scenario 1: Sample 10% of Lines (Entries)
If each line in your file counts as an "entry," this approach will randomly pick 10% of those lines without duplicates:
# 1. Read all lines from the original file (adjust encoding if needed) all_lines <- readLines("ABC.txt", encoding = "UTF-8") # 2. Calculate how many lines make up 10% of the total sample_size <- round(length(all_lines) * 0.1) # 3. Randomly select the lines sampled_lines <- sample(all_lines, size = sample_size, replace = FALSE) # 4. Save the sampled lines to a new file writeLines(sampled_lines, "sampled_ABC.txt", encoding = "UTF-8")
For Extra Large Files:
If your file is so big that loading all lines into memory at once causes issues, use a chunked approach to read and sample in batches:
# Open a connection to the file file_con <- file("ABC.txt", open = "r") sampled_lines <- c() chunk_size <- 10000 # Adjust based on your available memory # Read and sample in chunks until the end of the file while (length(chunk <- readLines(file_con, n = chunk_size)) > 0) { chunk_sample <- sample(chunk, size = round(length(chunk)*0.1), replace = FALSE) sampled_lines <- c(sampled_lines, chunk_sample) } # Close the connection and save the result close(file_con) writeLines(sampled_lines, "sampled_ABC.txt", encoding = "UTF-8")
Scenario 2: Sample 10% of Words
If your goal is to keep only 10% of all words in the file (like your example mentions), follow these steps:
# 1. Read all lines and combine into a single text string full_text <- paste(readLines("ABC.txt", encoding = "UTF-8"), collapse = " ") # 2. Optional: Clean up punctuation (adjust regex to match your needs) clean_text <- gsub("[^a-zA-Z0-9\\s]", "", full_text) # 3. Split the text into individual words all_words <- strsplit(clean_text, "\\s+")[[1]] # 4. Calculate 10% of the total word count sample_word_count <- round(length(all_words) * 0.1) # 5. Randomly select words and save to a new file sampled_words <- sample(all_words, size = sample_word_count, replace = FALSE) writeLines(paste(sampled_words, collapse = " "), "sampled_words_ABC.txt", encoding = "UTF-8")
A quick note: The punctuation cleaning step uses a regex that removes anything that isn't letters, numbers, or spaces. Feel free to tweak that regex if you need to keep specific punctuation (like hyphens in words).
内容的提问来源于stack exchange,提问作者user3443063

