使用awk将符合特定规则的第一列单词小写化的命令修正问题
Let's break down what's going wrong with your command and how to fix it.
Your Scenario
You want to convert the first column to lowercase only if it matches the regex ^[A-Z][a-z]+$, with fields separated by tabs.
- Input:
Agne XXX Agn_Ball YYY - Desired Output:
agne XXX Agn_Ball YYY - Original Command:
awk -F '\t' 'BEGIN {OFS='\t'} $1 ~ /^[A-Z][a-z]+$/ {print tolower($1),$2}1' < test.tmp > test.tmp2 - Actual (Incorrect) Output:
agne XXX Agne XXX Agn_Ball YYY
What's Causing the Duplicate Line?
The issue is the trailing 1 at the end of your awk script. In awk, any non-zero value (like 1) is treated as a true condition, which triggers the default action: print $0 (print the entire original line).
For lines where the first column matches your regex:
- Your script first runs
print tolower($1),$2(outputs the modified line) - Then the
1triggersprint $0(outputs the original unmodified line)
That's why you get duplicate lines for matching entries.
Fixed Commands
Here are two clean ways to fix this:
Option 1: Modify the Field and Print Once
This approach updates the first column in-place when it matches, then prints every line (modified or original) once:
awk -F '\t' 'BEGIN {OFS="\t"} $1 ~ /^[A-Z][a-z]+$/ { $1 = tolower($1) } 1' test.tmp > test.tmp2
- When the regex matches, we overwrite
$1with its lowercase version - The trailing
1then prints the entire line (with the modified$1if applicable)
Option 2: Explicit Conditional Printing
If you prefer separate logic for matching vs non-matching lines, use an if/else block (no trailing 1 needed):
awk -F '\t' 'BEGIN {OFS="\t"} { if ($1 ~ /^[A-Z][a-z]+$/) { print tolower($1), $2 } else { print $0 } }' test.tmp > test.tmp2
This explicitly handles each case, ensuring every line is printed exactly once.
Either of these will give you the desired output without duplicate lines.
内容的提问来源于stack exchange,提问作者Luca

