You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用grep/sed替换分号但保留引号内分号?(大文件场景)

高效替换大文件中分号(排除双引号内内容)的解决方案

Alright, let's tackle this problem head-on—you need to replace every semicolon with a comma in huge (up to 99GB) files, but leave semicolons inside double quotes untouched. The naive sed 's/;/,/g' won't cut it because it blasts through all semicolons, including those quoted. Here's how to do it efficiently without choking on memory:

方案1:Awk(推荐,可靠且高性能)

Awk is perfect for this scenario because it processes files line-by-line (no full-file memory load) and makes state tracking trivial. This script tracks whether we're inside a double quote and only replaces semicolons when we're outside:

awk -v quote_char='"' 'BEGIN{FS=""; OFS=""} {
    in_quote = 0
    for (i=1; i<=NF; i++) {
        # Toggle quote state when we hit a double quote
        if ($i == quote_char) in_quote = !in_quote
        # Replace semicolon only if we’re NOT inside quotes
        if (!in_quote && $i == ";") $i = ","
    }
    print $0
}' input.txt > output.txt

脚本解释:

  • FS=""; OFS="": Splits each line into individual characters and ensures we don't add extra spacing when rebuilding the line.
  • in_quote: A boolean flag that flips between 0 (outside quotes) and 1 (inside quotes) every time we hit a double quote.
  • The loop processes each character: only semicolons outside quotes get replaced with commas.
  • Fully stream-based: each line is processed and immediately written to output, so memory usage stays minimal even for 99GB files.

方案2:Sed(适合Sed爱好者)

If you prefer using Sed, this script uses label-based looping to handle multiple quote pairs per line, skipping semicolons inside quotes:

sed -e ':process_line' \
    -e '/"/!s/;/,/g; b' \
    -e 's/\([^"]*\)"\([^";]*\)\([";]\)/\1"\2\3/g' \
    -e 't process_line' input.txt > output.txt

脚本解释:

  • :process_line: A label to loop back to until all quote pairs are handled.
  • /"/!s/;/,/g; b: If the line has no quotes, replace all semicolons and exit the loop.
  • The regex matches content before a quote, content inside the quote (excluding semicolons/quotes), and the closing quote or semicolon—this preserves semicolons inside quotes.
  • t process_line: Jumps back to the label if a substitution was made, ensuring all quote pairs are processed.

性能优化小贴士

Since you're dealing with massive files, these tweaks will help speed things up:

  • Use GNU Awk (gawk) or GNU Sed: They're significantly faster than non-GNU implementations for large datasets.
  • Avoid tools that load the entire file into memory (e.g., Perl with -0777, Python's read()): They'll crash or crawl with 99GB files.
  • Use separate storage for input/output: If your input is on an SSD, write the output to a different disk/partition to avoid IO bottlenecks.
  • Monitor progress with pv: Pipe your input through pv to see real-time transfer rates and completion:
    pv input.txt | awk -v quote_char='"' 'BEGIN{FS=""; OFS=""} { ... }' > output.txt
    

测试验证

输入示例:
A;B;C;D 5cc0714b9b69581f14f6427f;5cc0714b9b69581f14f6428e;1;"5cc0714b9b69581f14f6427f;16a4fba8d13";xpto;

输出示例:
A,B,C,D 5cc0714b9b69581f14f6427f,5cc0714b9b69581f14f6428e,1,"5cc0714b9b69581f14f6427f;16a4fba8d13",xpto,

内容的提问来源于stack exchange,提问作者JonyD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:22:25