使用egrep统计多文件非表头行数遇异常,求EGREP/AWK/SED解决方案
Alright, let's fix this line-counting problem properly. The core issue with your original egrep command is that it's matching patterns anywhere in the line (not just header lines) and using an unescaped dot (which matches any character) for F220./Order., leading to false positives. Here are reliable solutions using egrep, awk, and sed that handle all your file cases correctly:
Solution 1: Fixed egrep Command
This version targets line-starting header patterns only, and escapes dots to match literal periods:
srccount=$(egrep -cv '^[#$]|^georgia:|^Order\.|^F220\.' "$FILE")
^anchors each pattern to the start of the line, so we only exclude lines that begin with your header markers.\.escapes the dot, soF220\.only matches lines starting withF220.(notF220xor other variations).-vinverts the match (count lines that don't match header patterns),-ccounts those lines.
This fixes both your edge cases:
- For files with
F220.headers: No more false positives on data lines withF220followed by non-dot characters. - For File4 (with
#/$headers): Only lines starting with#/$are excluded, so data lines containing those characters (if any) aren't incorrectly counted out.
Solution 2: awk (More Readable & Flexible)
awk makes it easy to explicitly define which lines to exclude, and avoids regex pitfalls:
srccount=$(awk '!/^[#$]/ && !/^georgia:/ && !/^Order\./ && !/^F220\./ {count++} END {print count+0}' "$FILE")
- The condition
!/pattern/skips lines matching each header pattern (again, anchored to line start with^). - We increment
countfor every non-header line, then print the total at the end.count+0ensures we output0if all lines are headers (instead of a blank value).
You can also write this more concisely with a single regex group:
srccount=$(awk '!/^([#$]|georgia:|Order\.|F220\.)/ {n++} END {print n+0}' "$FILE")
Solution 3: sed Alternative
If you prefer sed, this command deletes header lines then counts the remaining ones:
srccount=$(sed -n '/^[#$]/d; /^georgia:/d; /^Order\./d; /^F220\./d; =' "$FILE" | wc -l)
-nsuppresses default line printing./pattern/ddeletes each header line.=prints the line number for every remaining (non-header) line.wc -lcounts those line numbers to get the total non-header count.
All three solutions will correctly handle files with 1, 2, or 3 repeated headers, and won't break on any of your example file types.
内容的提问来源于stack exchange,提问作者Abhi

