大文件随机下采样:快速提取近似500万行样本需求
Awesome question! Given your file's uniform line lengths and that you just need "good enough" randomness (not mathematically perfect), we can leverage those details to sample blazingly fast—no need to load the entire 500MB file into memory at all. Below are the best methods tailored to your needs:
Method 1: Bash + awk (The Fastest Option)
Since your target sample size is exactly 5% of the total rows (5M / 100M), a probabilistic streaming approach with awk is unbeatable for speed. It processes the file in a single pass, no memory bloat, just raw disk I/O speed.
Run this command in your terminal:
awk 'BEGIN{srand()} rand() <= 0.05' input_file.txt > sampled_output.txt
How this works:
srand()seeds the random number generator once at the start.- For every line,
rand()spits out a value between 0 and 1. We keep the line if it’s ≤ 0.05 (a 5% chance). - Because your lines are all roughly the same length, the selected lines will be evenly spread across the entire file—perfect for "good enough" randomness.
- If you notice you’re consistently getting a few thousand lines less than 5M, tweak the threshold slightly (like
0.0505) to compensate for variance.
Method 2: Python (For Programmatic Workflows)
If you need to wrap this into a Python script, use the same streaming probabilistic approach to keep memory usage ultra-low:
import random import sys # 5% sampling probability (matches 5M/100M) SAMPLING_PROB = 0.05 def sample_large_file(input_path, output_path): random.seed() # Seed once for consistent randomness (remove for new samples each run) with open(input_path, 'r') as infile, open(output_path, 'w') as outfile: for line in infile: if random.random() <= SAMPLING_PROB: outfile.write(line) if __name__ == "__main__": if len(sys.argv) != 3: print("Usage: python sampler.py <input_file_path> <output_file_path>") sys.exit(1) sample_large_file(sys.argv[1], sys.argv[2])
Quick optimizations:
- This only holds one line in memory at a time, so your RAM won’t even break a sweat.
- For a tiny speed boost, add a larger buffer size to the
open()calls, likebuffering=1024*1024(1MB buffer).
Method 3: Precise Line Selection (If You Want Closer to Exact 5M Lines)
If the probabilistic variance bugs you, you can use your uniform line length to jump directly to random positions in the file. This avoids reading the whole file and gives you a more precise sample size:
import os import random def sample_precise_lines(input_path, output_path, target_lines=5000000): file_size = os.path.getsize(input_path) avg_line_size = file_size / 100000000 # Your known total row count max_offset = file_size - avg_line_size # Avoid hitting the end of the file # Generate random offsets and sort them (sequential reads are faster than random jumps) offsets = [random.uniform(0, max_offset) for _ in range(target_lines)] offsets.sort() with open(input_path, 'rb') as infile, open(output_path, 'wb') as outfile: for offset in offsets: infile.seek(int(offset)) infile.readline() # Skip to the start of the next full line line = infile.readline() if line: outfile.write(line) if __name__ == "__main__": sample_precise_lines("input_file.txt", "sampled_output.txt")
Why this works:
- Sorting the offsets means we read the file sequentially, which is way faster for disk drives than jumping around randomly.
- Since your lines are uniform, the random offsets map to roughly random lines across the entire file—no clustering.
Final Tips:
- Stick with the
awkmethod if you can—it’s the fastest by a mile, no setup required. - All these methods avoid loading the entire file into memory, which is critical for speed and keeping your system responsive.
- The probabilistic approaches give exactly the "good enough" randomness you’re looking for, without overcomplicating things.
内容的提问来源于stack exchange,提问作者ShaharA

