You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中按行顺序拆分1000万行Tab分隔文本数据集?

Got it, let's work through this dataset splitting problem. Since you have a 10M-row, tab-separated dataset and need to split it without shuffling rows or columns (keeping strict sequential order), here are two practical, reliable approaches depending on your workflow:

Approach 1: Python (Pandas for ML Workflows)

If you plan to do further machine learning preprocessing in Python, this method lets you split the file while keeping it compatible with libraries like scikit-learn or TensorFlow. We'll use chunked reading to avoid loading the entire 10M rows into memory at once.

First, calculate the row counts for each split (we'll assume you want to split the remaining 67% into two roughly equal parts—adjust the ratios as needed):

  • Total rows: 10,000,000
  • First part: 33% → 3,300,000 rows
  • Second part: ~33.5% → 3,350,000 rows
  • Third part: ~33.5% → 3,350,000 rows

Here's the code:

import pandas as pd

# Update these paths to match your files
input_path = "your_dataset.tsv"
train_path = "train_set.tsv"
val_path = "val_set.tsv"
test_path = "test_set.tsv"

# Get total row count and header
total_rows = sum(1 for _ in open(input_path, 'r'))
header = pd.read_csv(input_path, sep='\t', nrows=0).columns.tolist()

# Calculate split boundaries
first_split = int(total_rows * 0.33)
remaining = total_rows - first_split
second_split = first_split + (remaining // 2)

# Helper function to write a range of rows to a file
def write_rows(start_row, end_row, output_file):
    with open(output_file, 'w') as f_out:
        # Write header first
        f_out.write('\t'.join(header) + '\n')
        # Read in chunks to save memory
        chunk_size = 100000  # Adjust based on your available RAM
        for i, chunk in enumerate(pd.read_csv(input_path, sep='\t', chunksize=chunk_size, skiprows=1)):
            chunk_start = i * chunk_size
            chunk_end = chunk_start + chunk_size
            # Check if current chunk overlaps with our target range
            if chunk_end <= start_row:
                continue
            if chunk_start >= end_row:
                break
            # Slice the chunk to the target rows
            slice_start = max(0, start_row - chunk_start)
            slice_end = min(len(chunk), end_row - chunk_start)
            chunk.iloc[slice_start:slice_end].to_csv(f_out, sep='\t', index=False, header=False)

# Write each split
write_rows(0, first_split, train_path)
write_rows(first_split, second_split, val_path)
write_rows(second_split, total_rows, test_path)

Notes for this approach:

  • Adjust chunk_size based on how much RAM you have (smaller chunks use less memory but take slightly longer)
  • The write_rows function ensures we only write the exact rows we need, no extra data loaded
  • All split files retain the original header and column order
Approach 2: Command Line (Ultra-Fast for Large Files)

If you don't need Python for preprocessing, the command line is the fastest way to split your dataset—no need to load anything into memory, perfect for 10M-row files. This works on Linux/macOS (use Git Bash or WSL on Windows).

  1. First, get the total row count and header:
total_rows=$(wc -l < your_dataset.tsv)
header=$(head -n 1 your_dataset.tsv)
  1. Calculate split sizes:
first_part=$(( total_rows * 33 / 100 ))
remaining=$(( total_rows - first_part ))
second_part=$(( remaining / 2 ))
  1. Split the files (each gets the original header):
# Write first 33% (header + rows 2 to first_part+1)
echo "$header" > train_set.tsv
tail -n +2 your_dataset.tsv | head -n $first_part >> train_set.tsv

# Write second split
echo "$header" > val_set.tsv
tail -n +$(( first_part + 2 )) your_dataset.tsv | head -n $second_part >> val_set.tsv

# Write third split (remaining rows)
echo "$header" > test_set.tsv
tail -n +$(( first_part + second_part + 2 )) your_dataset.tsv >> test_set.tsv

Notes for this approach:

  • This is way faster than Python for pure file splitting—no overhead of data processing libraries
  • The math uses integer division, so the last split will have 1 extra row if remaining is odd (adjust second_part if you need exact equality)
  • All splits maintain strict row order, no shuffling at all

Quick Adjustment Tip:

If you have specific ratios for the second and third parts (not just splitting the remainder evenly), just update the second_part calculation to match your needs (e.g., second_part=$(( total_rows * 30 / 100 )) for a 33/30/37 split).

内容的提问来源于stack exchange,提问作者Amir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:47:10