You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:编写Bash循环批量拆分百万行单列文件为每4行一个文件

高效分割单列数据文件为每4行一组的Bash方案

Hey there! Let's get this sorted for you. You've got a single-column random number file table2.txt (around 1 million lines) and need to split it into chunks of 4 lines each—with the first chunk (lines 1-4) saved to file0.txt, the second (lines 5-8) to file1.txt, and so on, running this for 10001 iterations (since your loop starts at 0 and goes to 10000).

方案1:用split命令(最简洁高效)

split is built exactly for this kind of file splitting task, and it's way faster than looping with sed for large files. Here's the command you need:

split -l 4 --numeric-suffixes=0 --suffix-length=5 --additional-suffix=.txt table2.txt file

Let's break down the flags:

  • -l 4: Split the file into chunks of 4 lines each
  • --numeric-suffixes=0: Start the file suffixes at 0 instead of the default alphabetical ones
  • --suffix-length=5: Use 5-digit suffixes (so you get file00000.txt up to file10000.txt which matches your loop range)
  • --additional-suffix=.txt: Add the .txt extension to each output file
  • table2.txt: Your input file
  • file: The prefix for all output files

方案2:Bash循环 + sed(适合自定义逻辑)

If you need to stick with a bash loop for some reason, this will work—though note it's slower for large files since sed reads the entire file every time:

#!/bin/bash
for i in {0..10000}; do
    # Calculate start and end line numbers for each chunk
    start=$((i * 4 + 1))
    end=$((start + 3))
    # Extract the lines and save to the corresponding file
    sed -n "${start},${end}p" table2.txt > "file${i}.txt"
done

方案3:用awk(一次遍历,高效处理大文件)

awk is perfect for this because it only reads your input file once, making it way faster than the bash+sed approach for 1 million lines:

awk '{
    # Calculate which output file this line belongs to (starts at 0)
    file_num = int((NR - 1) / 4)
    # Write the line to the corresponding file
    print > "file" file_num ".txt"
}' table2.txt

This will automatically split your file into file0.txt, file1.txt, ..., all the way up to the last chunk—no need to hardcode the loop count if you don't want to.

注意事项

  • If your "single-column" file actually has multiple numbers per line (like your example shows), just adjust the logic to count items instead of lines—but your description says "单列随机数文件", so I assumed one number per line.
  • Make sure your input file has at least 40004 lines (4 lines × 10001 iterations) to avoid empty files at the end.

内容的提问来源于stack exchange,提问作者GB44444

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:41:07