如何用正则表达式移除800万行XML文件中的重复空sector标签?
Got it, let's solve this problem efficiently—dealing with an 8M-line XML file means we need solutions that don't choke on large data, while precisely targeting only duplicate empty <sector> tags (leaving single empty ones and all non-empty tags intact).
Regex Pattern & Replacement Logic
First, let's define what an empty <sector> tag looks like from your example: it has attributes, opening tag, whitespace-only content, and closing tag. We need to match consecutive duplicates of this pattern and replace them with just one instance.
Regex to Match Duplicate Empty Tags
Use this pattern for find-and-replace:
(<sector\s+[^>]+>\s*</sector>)(\s*\1)+
Replace with:
\1
Breakdown of the Pattern
(<sector\s+[^>]+>\s*</sector>): Capture group 1 matches a single empty<sector>tag:<sector\s+: Matches the opening tag start plus any leading whitespace before attributes[^>]+: Matches all attribute content (stops at the closing>of the opening tag)\s*: Matches any whitespace (spaces, newlines) inside the tag (since your example has a space, but large files might have line breaks here)</sector>: Matches the closing tag
(\s*\1)+: Matches one or more repetitions of the captured empty tag, allowing for whitespace (like newlines or spaces) between duplicates
Tools to Handle Large Files
GUI text editors might crash or lag with 8M lines—stick to command-line tools for efficiency:
Option 1: GNU Sed (Fast, Simple)
Run this command in your terminal:
sed -E 's/(<sector\s+[^>]+>\s*<\/sector>)(\s*\1)+/\1/g' input.xml > output.xml
-E: Enables extended regular expressions (avoids escaping parentheses)g: Ensures all duplicate instances in the file are replaced
Option 2: Perl (More Robust for Edge Cases)
Perl handles large files smoothly, and can also handle edge cases like empty tags split across lines. For global replacement:
perl -0777 -pe 's/(<sector\s+[^>]+>\s*<\/sector>)(\s*\1)+/\1/g' input.xml > output.xml
-0777: Reads the entire file at once (works if your system has enough RAM; if not, use the line-by-line script below)
If memory is a concern, use this Perl script to process line-by-line without loading the whole file:
#!/usr/bin/perl use strict; use warnings; my $last_empty_tag = ''; while (<>) { chomp; # Check if current line is an empty sector tag if (/^(<sector\s+[^>]+>\s*<\/sector>)$/) { # Only print if it's different from the last empty tag we saw if ($1 ne $last_empty_tag) { print "$1\n"; $last_empty_tag = $1; } } else { # Print non-empty tags or other content, reset the last empty tag tracker print "$_\n"; $last_empty_tag = ''; } }
Save this as remove_duplicates.pl, make it executable (chmod +x remove_duplicates.pl), then run:
./remove_duplicates.pl input.xml > output.xml
Important Notes
- Backup First: Always make a copy of your original XML file before running any replacements—better safe than sorry!
- Tag Format Consistency: This regex assumes your empty
<sector>tags have attributes on the same line as the opening tag, and no nested tags inside. If your XML has more complex formatting (e.g., attributes split across lines), adjust the pattern to(<sector\s+[\s\S]*?>\s*</sector>)(\s*\1)+(use non-greedy matching for attributes). - XML vs Regex Caveat: Regex isn't a full XML parser, but since we're targeting flat, non-nested
<sector>tags, this approach is safe and efficient for your use case.
内容的提问来源于stack exchange,提问作者RogerHN

