如何用grep/sed替换分号但保留引号内分号?(大文件场景)
Alright, let's tackle this problem head-on—you need to replace every semicolon with a comma in huge (up to 99GB) files, but leave semicolons inside double quotes untouched. The naive sed 's/;/,/g' won't cut it because it blasts through all semicolons, including those quoted. Here's how to do it efficiently without choking on memory:
方案1:Awk(推荐,可靠且高性能)
Awk is perfect for this scenario because it processes files line-by-line (no full-file memory load) and makes state tracking trivial. This script tracks whether we're inside a double quote and only replaces semicolons when we're outside:
awk -v quote_char='"' 'BEGIN{FS=""; OFS=""} { in_quote = 0 for (i=1; i<=NF; i++) { # Toggle quote state when we hit a double quote if ($i == quote_char) in_quote = !in_quote # Replace semicolon only if we’re NOT inside quotes if (!in_quote && $i == ";") $i = "," } print $0 }' input.txt > output.txt
脚本解释:
FS=""; OFS="": Splits each line into individual characters and ensures we don't add extra spacing when rebuilding the line.in_quote: A boolean flag that flips between 0 (outside quotes) and 1 (inside quotes) every time we hit a double quote.- The loop processes each character: only semicolons outside quotes get replaced with commas.
- Fully stream-based: each line is processed and immediately written to output, so memory usage stays minimal even for 99GB files.
方案2:Sed(适合Sed爱好者)
If you prefer using Sed, this script uses label-based looping to handle multiple quote pairs per line, skipping semicolons inside quotes:
sed -e ':process_line' \ -e '/"/!s/;/,/g; b' \ -e 's/\([^"]*\)"\([^";]*\)\([";]\)/\1"\2\3/g' \ -e 't process_line' input.txt > output.txt
脚本解释:
:process_line: A label to loop back to until all quote pairs are handled./"/!s/;/,/g; b: If the line has no quotes, replace all semicolons and exit the loop.- The regex matches content before a quote, content inside the quote (excluding semicolons/quotes), and the closing quote or semicolon—this preserves semicolons inside quotes.
t process_line: Jumps back to the label if a substitution was made, ensuring all quote pairs are processed.
性能优化小贴士
Since you're dealing with massive files, these tweaks will help speed things up:
- Use GNU Awk (gawk) or GNU Sed: They're significantly faster than non-GNU implementations for large datasets.
- Avoid tools that load the entire file into memory (e.g., Perl with
-0777, Python'sread()): They'll crash or crawl with 99GB files. - Use separate storage for input/output: If your input is on an SSD, write the output to a different disk/partition to avoid IO bottlenecks.
- Monitor progress with
pv: Pipe your input throughpvto see real-time transfer rates and completion:pv input.txt | awk -v quote_char='"' 'BEGIN{FS=""; OFS=""} { ... }' > output.txt
测试验证
输入示例:
A;B;C;D 5cc0714b9b69581f14f6427f;5cc0714b9b69581f14f6428e;1;"5cc0714b9b69581f14f6427f;16a4fba8d13";xpto;
输出示例:
A,B,C,D 5cc0714b9b69581f14f6427f,5cc0714b9b69581f14f6428e,1,"5cc0714b9b69581f14f6427f;16a4fba8d13",xpto,
内容的提问来源于stack exchange,提问作者JonyD

