使用Bash处理CSV文件:移除首尾引号及字段首尾双引号
Got it, let's fix those extra double quotes in your CSV file. The problem breaks down into two key steps: removing the wrapping quotes around each entire line, then fixing the escaped quotes around individual fields. Here are two straightforward ways to do this with standard Bash tools:
1. Using sed (Stream Editor)
sed is perfect for quick text substitutions. We'll chain three replacement commands to handle the job:
sed -e 's/^"//' -e 's/"$//' -e 's/""/"/g' your_input.csv > cleaned_output.csv
Let's break down what each part does:
s/^"//: Removes the double quote at the start of each lines/"$//: Removes the double quote at the end of each lines/""/"/g: Globally replaces every instance of two consecutive double quotes ("") with a single double quote ("), which fixes the escaped field quotes
2. Using awk
If you prefer awk (great for more structured text processing), this command achieves the same result:
awk '{ # Remove leading and trailing quotes from the line gsub(/^"|"$/, ""); # Replace all escaped double quotes with single quotes gsub(/""/, "\""); # Print the cleaned line print }' your_input.csv > cleaned_output.csv
Example Test
For your input line:
"2,""Peter"",""Parker"",""30"",""NY"""
Both commands will output:2,"Peter","Parker","30","NY"
Note
If your CSV has more complex edge cases (like fields containing commas or actual double quotes as part of the data), you might want to use a dedicated CSV processing tool like csvkit to avoid breaking the structure. But for the specific format you described, the above commands work perfectly.
内容的提问来源于stack exchange,提问作者Mangekyou

