Linux中如何用grep或其他工具跨行匹配指定多行文本模式?
Hey there! I see you're a Linux newbie trying to match text patterns that span multiple lines—specifically content starting with "Leonard is" and ending with "champion" (case-insensitive). Your original grep command only catches single-line matches because grep processes lines one by default, so .* can't cross line breaks, and $ only matches the end of a single line. Let's fix this!
问题分析
Your input file has lines where the target pattern wraps across lines (like line 3's "Leonard is a" and line 4's "Champion"). The default grep behavior can't handle this, so we need tools or options that support multi-line matching, plus we want to keep the original line numbers as shown in your desired output.
解决方案1:使用Awk(最贴合你的期望输出)
Awk is perfect here because we can track when we start matching "Leonard is" and keep collecting content until we hit "champion", while preserving line numbers.
Create a script file (e.g., match_pattern.awk) with this content:
BEGIN { IGNORECASE = 1 # 启用不区分大小写匹配 in_match = 0 # 标记是否处于跨行匹配状态 buffer = "" # 存储跨行匹配的内容 } { # 提取每行的行号和内容 num = substr($0, 1, index($0, " ") - 1) content = substr($0, index($0, " ") + 1) if (in_match) { # 将当前行的行号和内容加入缓冲区 buffer = buffer " " num " " content # 检查是否到达模式结尾(champion) if (content ~ /champion$/) { print buffer in_match = 0 buffer = "" } } else { # 检查单行内是否有完整匹配 if (match(content, /Leonard is.*champion$/, match_arr)) { print num " " match_arr[0] } # 检查是否开始跨行匹配 else if (content ~ /Leonard is/) { in_match = 1 buffer = num " " content } } }
运行命令:
awk -f match_pattern.awk input.txt
这会输出你期望的结果:
1 Leonard is a champion 3 Leonard is a 4 Champion 5 Leonard is An exemplary 6 Champion
解决方案2:使用pcregrep(快速多行匹配,不保留行号格式)
如果不需要保留跨行后的行号格式,pcregrep(支持PCRE多行正则)是更简单的选择。先安装它(比如Debian/Ubuntu系统:sudo apt install pcregrep):
pcregrep -ioM '\bLeonard is\b.*?\bchampion\b' input.txt
-M: 启用多行匹配-i: 不区分大小写-o: 只输出匹配的部分.*?: 非贪婪匹配,避免一次性捕获多个"champion"实例\b: 单词边界,防止部分匹配(比如"Leonard isss")
输出结果:
Leonard is a champion Leonard is a Champion Leonard is An exemplary Champion
解决方案3:Tr + Grep(简单但丢失行号)
如果只需要匹配到的文本、不需要行号,可以先把所有换行转为空格,再用grep匹配:
tr '\n' ' ' < input.txt | grep -ioE 'Leonard is.*?champion'
这个方法把整个文件转为单行,让grep的.*?可以跨越原本的换行符进行匹配。
内容的提问来源于stack exchange,提问作者Hammad Ahmed

