如何将9915万条字符数据按1000行分块写入本地文件?
Chunking a Large Character Vector into CSV Files (1000 Rows Each)
Got it, let's adapt your existing dataframe chunking logic to work with a large character vector—this is straightforward and will be efficient for your 99M-row dataset.
Since your character vector already contains fully formatted CSV rows (each element is a comma-separated string), we can skip converting it to a dataframe entirely. Instead, we'll directly slice the vector into chunks and write each chunk to a file using writeLines() (which is faster for raw character data than write.table()).
Here's the adapted code:
# Define your chunk size and target directory chunk_size <- 1000 output_dir <- "path_to_local_dir/" # Replace char_data with the name of your actual character vector total_rows <- length(char_data) # Iterate over each chunk start index lapply(seq(1, total_rows, by = chunk_size), function(i) { # Calculate the end index for the current chunk (handles the final partial chunk) end_idx <- min(i + chunk_size - 1, total_rows) # Extract the chunk of character rows current_chunk <- char_data[i:end_idx] # Create a unique filename matching your original pattern file_path <- paste0(output_dir, "data", i, "-", end_idx, ".csv") # Write the chunk to file writeLines(current_chunk, con = file_path) })
Key details to note:
writeLines()is ideal here: It writes each element of the character vector as a separate line in the file, which is exactly what we need for CSV (since each element is already a properly formatted row).- Handling the final chunk: The
min(i + chunk_size - 1, total_rows)ensures we don't try to index beyond the length of your vector for the last chunk (which will have 500 rows in your case, since 99150000 = 99150*1000). - Efficiency: By working directly with the character vector, we avoid the memory overhead of converting to a dataframe first—critical for such a large dataset.
If you need to specify a file encoding (e.g., UTF-8), just add the encoding parameter to writeLines():
writeLines(current_chunk, con = file_path, encoding = "UTF-8")
内容的提问来源于stack exchange,提问作者khajlk
相关产品推荐
相关产品推荐

