You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用sed+awk按字段数分类行至文件时的非原行输出问题

Hey, let's break down why your command isn't writing the original lines, and fix it properly!

The Root Cause

Your current pipeline first uses sed to strip off trailing ||o|| sequences from every line before passing it to awk. That means the lines awk writes to 1.txt and not_1.txt are already modified by sed—not the original lines from your sample file. That's exactly why you're seeing unexpected output.

The Fix: Preserve Original Lines with Awk

We can handle everything in a single awk command, so we keep the original line content while still correctly categorizing lines by field count:

awk '
BEGIN { FS = /\|\|o\|\|/ }
{
    # Save the original unmodified line
    original_line = $0
    # Make a temp copy to manipulate for field counting
    temp_line = $0
    # Remove trailing ||o|| sequences from the temp line only
    sub(/(\|\|o\|\|)+$/, "", temp_line)
    # Split the cleaned temp line to get actual non-empty field count
    field_count = split(temp_line, _, FS)
    
    # Write the original line to the correct file
    if (field_count == 1) {
        print original_line > "1.txt"
        close("1.txt")  # Optional but recommended for large files
    } else {
        print original_line > "not_1.txt"
        close("not_1.txt")
    }
}
' sample.txt

Code Breakdown

  • BEGIN { FS = /\|\|o\|\|/ }: Sets the field separator to ||o||. We escape the | characters because they're special in regular expressions.
  • original_line = $0: Stores the unmodified current line in a variable—this is what we'll write to the output files, ensuring we keep your original content.
  • sub(/(\|\|o\|\|)+$/, "", temp_line): Removes only trailing ||o|| sequences from the temporary copy of the line, so we don't count empty fields at the end when determining field count.
  • split(temp_line, _, FS): Splits the cleaned temp line to get the actual number of meaningful fields (ignoring trailing empty ones).
  • close(): If you're working with very large files, closing the files after writing prevents awk from running out of file handles. For small files, you can safely skip this step.

With this command, 1.txt will contain your original single-field lines, and not_1.txt will hold the original multi-field lines—exactly what you wanted!

内容的提问来源于stack exchange,提问作者Bhawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:39:24