关于sed命令的技术咨询:筛选CSV文件特定条件行
Hey there! Let's work through this CSV filtering issue. First off, I want to note that sed is great for text pattern matching, but it doesn't handle numeric comparisons natively—that's probably why your earlier attempts were tripping up, especially with Val2's range. awk is the far more reliable tool here because it can directly work with numeric values, even when your decimal separator is a comma. Let's cover both approaches, starting with the recommended one.
Recommended: Use awk for Numeric Filtering
Assuming your CSV uses commas as field separators, and Val1 is in column 1, Val2 in column 2 (adjust the field numbers if yours are different), here's the command:
-F',' 'BEGIN {OFS=","} { # Replace comma with dot in Val2 to treat it as a numeric value gsub(/,/, ".", $2) # Apply your filtering conditions if ($1 > 3500 && $2 >= 60.00 && $2 <= 99.99) { print $0 } }' your_file.csv
Breakdown:
-F',': Tellsawkto split each line on commas (adjust if your CSV uses a different separator like tabs).BEGIN {OFS=","}: Ensures the output CSV retains commas as field separators, so your output stays properly formatted.gsub(/,/, ".", $2): Converts the comma in Val2 to a dot—this letsawktreat the value as a number instead of a string, making the range check accurate.- The
ifstatement enforces your two conditions:- Val1 (column 1) is strictly greater than 3500
- Val2 (column 2) is between 60.00 and 99.99 (inclusive)
print $0: Outputs the entire line if it passes both checks.
Skip Header Row (If You Have One):
Add NR>1 to skip the first line (your header):
-F',' 'BEGIN {OFS=","} NR>1 { gsub(/,/, ".", $2) if ($1 > 3500 && $2 >= 60.00 && $2 <= 99.99) { print $0 } }' your_file.csv
If You Must Use sed (Text Pattern Matching Only)
Since sed can't do true numeric comparisons, we have to use regex to match the string patterns that correspond to your numeric ranges. This is less robust (e.g., it fails if Val1 has leading zeros), but here's how to do it:
Regex Breakdown:
- Val1 > 3500: Matches strings that are either:
- 5+ digits (
[0-9]\{5,\}), - 4-digit numbers starting with 36-99 (
[3-9][6-9][0-9]\{2\}), - 4-digit numbers starting with 3501-3599 (
[3][5][0-9][1-9]).
- 5+ digits (
- Val2 60,00-99,99: Matches strings that are:
- 60-69 followed by
,and two digits (6[0-9],[0-9]\{2\}), - 70-89 followed by
,and two digits ([7-8][0-9],[0-9]\{2\}), - 90-99 followed by
,and two digits (9[0-9],[0-9]\{2\}).
- 60-69 followed by
Again, this regex approach is fragile—if your data has any edge cases (like leading zeros, spaces in fields, or Val1 values like 3500.5), it won't work correctly. Stick with awk for reliability.
内容的提问来源于stack exchange,提问作者antocad

