使用sed+awk按字段数分类行至文件时的非原行输出问题
Hey, let's break down why your command isn't writing the original lines, and fix it properly!
The Root Cause
Your current pipeline first uses sed to strip off trailing ||o|| sequences from every line before passing it to awk. That means the lines awk writes to 1.txt and not_1.txt are already modified by sed—not the original lines from your sample file. That's exactly why you're seeing unexpected output.
The Fix: Preserve Original Lines with Awk
We can handle everything in a single awk command, so we keep the original line content while still correctly categorizing lines by field count:
awk ' BEGIN { FS = /\|\|o\|\|/ } { # Save the original unmodified line original_line = $0 # Make a temp copy to manipulate for field counting temp_line = $0 # Remove trailing ||o|| sequences from the temp line only sub(/(\|\|o\|\|)+$/, "", temp_line) # Split the cleaned temp line to get actual non-empty field count field_count = split(temp_line, _, FS) # Write the original line to the correct file if (field_count == 1) { print original_line > "1.txt" close("1.txt") # Optional but recommended for large files } else { print original_line > "not_1.txt" close("not_1.txt") } } ' sample.txt
Code Breakdown
BEGIN { FS = /\|\|o\|\|/ }: Sets the field separator to||o||. We escape the|characters because they're special in regular expressions.original_line = $0: Stores the unmodified current line in a variable—this is what we'll write to the output files, ensuring we keep your original content.sub(/(\|\|o\|\|)+$/, "", temp_line): Removes only trailing||o||sequences from the temporary copy of the line, so we don't count empty fields at the end when determining field count.split(temp_line, _, FS): Splits the cleaned temp line to get the actual number of meaningful fields (ignoring trailing empty ones).close(): If you're working with very large files, closing the files after writing prevents awk from running out of file handles. For small files, you can safely skip this step.
With this command, 1.txt will contain your original single-field lines, and not_1.txt will hold the original multi-field lines—exactly what you wanted!
内容的提问来源于stack exchange,提问作者Bhawan

