如何在Python中按行顺序拆分1000万行Tab分隔文本数据集?
Got it, let's work through this dataset splitting problem. Since you have a 10M-row, tab-separated dataset and need to split it without shuffling rows or columns (keeping strict sequential order), here are two practical, reliable approaches depending on your workflow:
If you plan to do further machine learning preprocessing in Python, this method lets you split the file while keeping it compatible with libraries like scikit-learn or TensorFlow. We'll use chunked reading to avoid loading the entire 10M rows into memory at once.
First, calculate the row counts for each split (we'll assume you want to split the remaining 67% into two roughly equal parts—adjust the ratios as needed):
- Total rows: 10,000,000
- First part: 33% → 3,300,000 rows
- Second part: ~33.5% → 3,350,000 rows
- Third part: ~33.5% → 3,350,000 rows
Here's the code:
import pandas as pd # Update these paths to match your files input_path = "your_dataset.tsv" train_path = "train_set.tsv" val_path = "val_set.tsv" test_path = "test_set.tsv" # Get total row count and header total_rows = sum(1 for _ in open(input_path, 'r')) header = pd.read_csv(input_path, sep='\t', nrows=0).columns.tolist() # Calculate split boundaries first_split = int(total_rows * 0.33) remaining = total_rows - first_split second_split = first_split + (remaining // 2) # Helper function to write a range of rows to a file def write_rows(start_row, end_row, output_file): with open(output_file, 'w') as f_out: # Write header first f_out.write('\t'.join(header) + '\n') # Read in chunks to save memory chunk_size = 100000 # Adjust based on your available RAM for i, chunk in enumerate(pd.read_csv(input_path, sep='\t', chunksize=chunk_size, skiprows=1)): chunk_start = i * chunk_size chunk_end = chunk_start + chunk_size # Check if current chunk overlaps with our target range if chunk_end <= start_row: continue if chunk_start >= end_row: break # Slice the chunk to the target rows slice_start = max(0, start_row - chunk_start) slice_end = min(len(chunk), end_row - chunk_start) chunk.iloc[slice_start:slice_end].to_csv(f_out, sep='\t', index=False, header=False) # Write each split write_rows(0, first_split, train_path) write_rows(first_split, second_split, val_path) write_rows(second_split, total_rows, test_path)
Notes for this approach:
- Adjust
chunk_sizebased on how much RAM you have (smaller chunks use less memory but take slightly longer) - The
write_rowsfunction ensures we only write the exact rows we need, no extra data loaded - All split files retain the original header and column order
If you don't need Python for preprocessing, the command line is the fastest way to split your dataset—no need to load anything into memory, perfect for 10M-row files. This works on Linux/macOS (use Git Bash or WSL on Windows).
- First, get the total row count and header:
total_rows=$(wc -l < your_dataset.tsv) header=$(head -n 1 your_dataset.tsv)
- Calculate split sizes:
first_part=$(( total_rows * 33 / 100 )) remaining=$(( total_rows - first_part )) second_part=$(( remaining / 2 ))
- Split the files (each gets the original header):
# Write first 33% (header + rows 2 to first_part+1) echo "$header" > train_set.tsv tail -n +2 your_dataset.tsv | head -n $first_part >> train_set.tsv # Write second split echo "$header" > val_set.tsv tail -n +$(( first_part + 2 )) your_dataset.tsv | head -n $second_part >> val_set.tsv # Write third split (remaining rows) echo "$header" > test_set.tsv tail -n +$(( first_part + second_part + 2 )) your_dataset.tsv >> test_set.tsv
Notes for this approach:
- This is way faster than Python for pure file splitting—no overhead of data processing libraries
- The math uses integer division, so the last split will have 1 extra row if
remainingis odd (adjustsecond_partif you need exact equality) - All splits maintain strict row order, no shuffling at all
Quick Adjustment Tip:
If you have specific ratios for the second and third parts (not just splitting the remainder evenly), just update the second_part calculation to match your needs (e.g., second_part=$(( total_rows * 30 / 100 )) for a 33/30/37 split).
内容的提问来源于stack exchange,提问作者Amir

