You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用egrep统计多文件非表头行数遇异常,求EGREP/AWK/SED解决方案

Alright, let's fix this line-counting problem properly. The core issue with your original egrep command is that it's matching patterns anywhere in the line (not just header lines) and using an unescaped dot (which matches any character) for F220./Order., leading to false positives. Here are reliable solutions using egrep, awk, and sed that handle all your file cases correctly:

Solution 1: Fixed egrep Command

This version targets line-starting header patterns only, and escapes dots to match literal periods:

srccount=$(egrep -cv '^[#$]|^georgia:|^Order\.|^F220\.' "$FILE")
  • ^ anchors each pattern to the start of the line, so we only exclude lines that begin with your header markers.
  • \. escapes the dot, so F220\. only matches lines starting with F220. (not F220x or other variations).
  • -v inverts the match (count lines that don't match header patterns), -c counts those lines.

This fixes both your edge cases:

  • For files with F220. headers: No more false positives on data lines with F220 followed by non-dot characters.
  • For File4 (with #/$ headers): Only lines starting with #/$ are excluded, so data lines containing those characters (if any) aren't incorrectly counted out.

Solution 2: awk (More Readable & Flexible)

awk makes it easy to explicitly define which lines to exclude, and avoids regex pitfalls:

srccount=$(awk '!/^[#$]/ && !/^georgia:/ && !/^Order\./ && !/^F220\./ {count++} END {print count+0}' "$FILE")
  • The condition !/pattern/ skips lines matching each header pattern (again, anchored to line start with ^).
  • We increment count for every non-header line, then print the total at the end. count+0 ensures we output 0 if all lines are headers (instead of a blank value).

You can also write this more concisely with a single regex group:

srccount=$(awk '!/^([#$]|georgia:|Order\.|F220\.)/ {n++} END {print n+0}' "$FILE")

Solution 3: sed Alternative

If you prefer sed, this command deletes header lines then counts the remaining ones:

srccount=$(sed -n '/^[#$]/d; /^georgia:/d; /^Order\./d; /^F220\./d; =' "$FILE" | wc -l)
  • -n suppresses default line printing.
  • /pattern/d deletes each header line.
  • = prints the line number for every remaining (non-header) line.
  • wc -l counts those line numbers to get the total non-header count.

All three solutions will correctly handle files with 1, 2, or 3 repeated headers, and won't break on any of your example file types.

内容的提问来源于stack exchange,提问作者Abhi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:31:25