You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linux中如何用grep或其他工具跨行匹配指定多行文本模式?

解决Linux中跨行匹配文本模式的问题

Hey there! I see you're a Linux newbie trying to match text patterns that span multiple lines—specifically content starting with "Leonard is" and ending with "champion" (case-insensitive). Your original grep command only catches single-line matches because grep processes lines one by default, so .* can't cross line breaks, and $ only matches the end of a single line. Let's fix this!

问题分析

Your input file has lines where the target pattern wraps across lines (like line 3's "Leonard is a" and line 4's "Champion"). The default grep behavior can't handle this, so we need tools or options that support multi-line matching, plus we want to keep the original line numbers as shown in your desired output.

解决方案1:使用Awk(最贴合你的期望输出)

Awk is perfect here because we can track when we start matching "Leonard is" and keep collecting content until we hit "champion", while preserving line numbers.

Create a script file (e.g., match_pattern.awk) with this content:

BEGIN { 
    IGNORECASE = 1  # 启用不区分大小写匹配
    in_match = 0     # 标记是否处于跨行匹配状态
    buffer = ""      # 存储跨行匹配的内容
}

{
    # 提取每行的行号和内容
    num = substr($0, 1, index($0, " ") - 1)
    content = substr($0, index($0, " ") + 1)

    if (in_match) {
        # 将当前行的行号和内容加入缓冲区
        buffer = buffer " " num " " content
        
        # 检查是否到达模式结尾(champion)
        if (content ~ /champion$/) {
            print buffer
            in_match = 0
            buffer = ""
        }
    } else {
        # 检查单行内是否有完整匹配
        if (match(content, /Leonard is.*champion$/, match_arr)) {
            print num " " match_arr[0]
        } 
        # 检查是否开始跨行匹配
        else if (content ~ /Leonard is/) {
            in_match = 1
            buffer = num " " content
        }
    }
}

运行命令:

awk -f match_pattern.awk input.txt

这会输出你期望的结果:

1 Leonard is a champion
3 Leonard is a 4 Champion
5 Leonard is An exemplary 6 Champion

解决方案2:使用pcregrep(快速多行匹配,不保留行号格式)

如果不需要保留跨行后的行号格式,pcregrep(支持PCRE多行正则)是更简单的选择。先安装它(比如Debian/Ubuntu系统:sudo apt install pcregrep):

pcregrep -ioM '\bLeonard is\b.*?\bchampion\b' input.txt
  • -M: 启用多行匹配
  • -i: 不区分大小写
  • -o: 只输出匹配的部分
  • .*?: 非贪婪匹配,避免一次性捕获多个"champion"实例
  • \b: 单词边界,防止部分匹配(比如"Leonard isss")

输出结果:

Leonard is a champion
Leonard is a Champion
Leonard is An exemplary Champion

解决方案3:Tr + Grep(简单但丢失行号)

如果只需要匹配到的文本、不需要行号,可以先把所有换行转为空格,再用grep匹配:

tr '\n' ' ' < input.txt | grep -ioE 'Leonard is.*?champion'

这个方法把整个文件转为单行,让grep的.*?可以跨越原本的换行符进行匹配。

内容的提问来源于stack exchange,提问作者Hammad Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:53:20