Linux下处理跨多行Windows CSV文件的引号转义与计数问题
Got it, let's tackle this CSV cleanup problem step by step. Your approach using a temporary placeholder (~) is smart—let's fix the kinks in your commands and get this working properly.
First, let's break down why your initial commands didn't work:
Why Your grep Command Failed
Your grep '"\r\n"\d{16}~' | wc -l command didn't work for three key reasons:
- Regex Compatibility:
grepdefaults to Basic Regular Expressions (BRE), which doesn't recognize\d(use[0-9]instead) or{16}without escaping. - Line-by-Line Processing:
grepreads input line by line, so it can't match the cross-line pattern"\r\n". - Windows Line Endings:
\r(carriage return) isn't always handled correctly in standardgrepunless you use Perl-compatible regex.
A working version to count those cross-line matches would be:
grep -zPo '"(\r?\n)"[0-9]{16}~' temp1.txt | wc -l
-z: Treats the entire file as a single "line" (ignores line breaks for matching)-P: Enables Perl-compatible regex (supports\rand clean quantifiers like{16})-o: Only outputs matched parts, sowc -lcounts correctly
Why Your sed Cross-Line Replace Failed
sed processes input line by line by default, so it can't see the "\r\n" spanning two lines. You need to force sed to read the entire file into memory first to handle cross-line patterns.
Full Working Solution
Let's walk through the corrected steps with proper commands, assuming your input file is named windows_csv.txt:
1. Replace | with ~|~ (Temporary Separator)
You already got this right, but let's formalize it:
sed 's/|/~|~/g' windows_csv.txt > temp1.txt
2. Replace Leading/Trailing Quotes in the File
Fix the first line's opening quote and last line's closing quote:
sed -i '1 s/^"/~/; $ s/"$/~/' temp1.txt
1 s/^"/~/: Targets the first line, replaces leading"with~$ s/"$/~: Targets the last line, replaces trailing"with~
3. Fix Cross-Line Field Boundaries
This is the tricky part—we need to replace the " at the end of a line and the " at the start of the next line (when followed by a 16-digit number and ~) with ~. Use sed in multi-line mode:
# For Windows line endings (\r\n) sed -i.bak ':loop; N; $!b loop; s/"\r\n"\([0-9]\{16\}~\)/~\r\n~\1/g' temp1.txt # If you converted to Linux line endings (\n) with dos2unix first: # sed -i.bak ':loop; N; $!b loop; s/"\n"\([0-9]\{16\}~\)/~\n~\1/g' temp1.txt
:loop; N; $!b loop: Reads the entire file into memory so we can match cross-line patternss/"\r\n"\([0-9]\{16\}~\)/~\r\n~\1/g: Replaces the cross-line quote pair with~while preserving the 16-digit field
4. Escape Embedded Quotes
Replace all remaining unescaped " with "" (the CSV standard for escaping):
sed -i 's/"/""/g' temp1.txt
5. Restore Original Quote Delimiters
Swap the temporary ~ back to ":
sed -i 's/~/""/g' temp1.txt
Alternative: Use Perl for Simpler Cross-Line Handling
Perl makes multi-line regex and CSV edge cases easier. Here's a one-liner that does all steps in one go:
perl -0777 -pe ' s/\|/~|~/g; s/"(\r?\n)"([0-9]{16}~)/~\1~\2/g; s/^"/~/; s/"$/~/"g; s/"/""/g; s/~/""/g; ' windows_csv.txt > fixed_csv.txt
-0777: Reads the entire file into memory (same as thesedloop trick)- The regexes handle Windows/Linux line endings (
\r?\n) automatically
内容的提问来源于stack exchange,提问作者always_learning

