You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言readlines从大文本文件随机提取10%条目

Hey there! Let's break down how to solve your problem—whether you're looking to sample 10% of lines (entries) or 10% of words from your large TXT file using R's readLines() function. Here are two scenarios tailored to your needs:

Scenario 1: Sample 10% of Lines (Entries)

If each line in your file counts as an "entry," this approach will randomly pick 10% of those lines without duplicates:

# 1. Read all lines from the original file (adjust encoding if needed)
all_lines <- readLines("ABC.txt", encoding = "UTF-8")

# 2. Calculate how many lines make up 10% of the total
sample_size <- round(length(all_lines) * 0.1)

# 3. Randomly select the lines
sampled_lines <- sample(all_lines, size = sample_size, replace = FALSE)

# 4. Save the sampled lines to a new file
writeLines(sampled_lines, "sampled_ABC.txt", encoding = "UTF-8")

For Extra Large Files:

If your file is so big that loading all lines into memory at once causes issues, use a chunked approach to read and sample in batches:

# Open a connection to the file
file_con <- file("ABC.txt", open = "r")
sampled_lines <- c()
chunk_size <- 10000  # Adjust based on your available memory

# Read and sample in chunks until the end of the file
while (length(chunk <- readLines(file_con, n = chunk_size)) > 0) {
  chunk_sample <- sample(chunk, size = round(length(chunk)*0.1), replace = FALSE)
  sampled_lines <- c(sampled_lines, chunk_sample)
}

# Close the connection and save the result
close(file_con)
writeLines(sampled_lines, "sampled_ABC.txt", encoding = "UTF-8")

Scenario 2: Sample 10% of Words

If your goal is to keep only 10% of all words in the file (like your example mentions), follow these steps:

# 1. Read all lines and combine into a single text string
full_text <- paste(readLines("ABC.txt", encoding = "UTF-8"), collapse = " ")

# 2. Optional: Clean up punctuation (adjust regex to match your needs)
clean_text <- gsub("[^a-zA-Z0-9\\s]", "", full_text)

# 3. Split the text into individual words
all_words <- strsplit(clean_text, "\\s+")[[1]]

# 4. Calculate 10% of the total word count
sample_word_count <- round(length(all_words) * 0.1)

# 5. Randomly select words and save to a new file
sampled_words <- sample(all_words, size = sample_word_count, replace = FALSE)
writeLines(paste(sampled_words, collapse = " "), "sampled_words_ABC.txt", encoding = "UTF-8")

A quick note: The punctuation cleaning step uses a regex that removes anything that isn't letters, numbers, or spaces. Feel free to tweak that regex if you need to keep specific punctuation (like hyphens in words).

内容的提问来源于stack exchange,提问作者user3443063

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:05:26