You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python3中打印跨两行的正则表达式匹配结果

问题:Python正则仅捕获分隔符行,无法匹配后续书签行

我写了个Python3脚本读取书签文件,用的正则在Notepad++里能正常匹配分隔符==========和后续的书签行(比如Book1 (Author 1)),但在Python里只能捕获分隔符行,拿不到后面的书签内容。

文本文件内容:

==========
Book1 (Author 1)
- bookmark

text
==========
Book2 (Author 2)
- bookmark1

text
==========
Book1 (Author 1)
- bookmark2

text
==========
Book2 (Author 2)
- bookmark2

text
==========

原Python脚本:

import re
pattern = re.compile("(==========)([\r\n])(.*)")
count=0
for line in open(r'bookmarks.txt', encoding="utf-8"):
    for match in re.finditer(pattern, line):
        count=count+1
        print(line)
print("The amount of notes are: ",count)

问题原因:

  1. 逐行读取的限制:原脚本逐行读取文件,分隔符行和书签行是独立的两行,正则没法在单行里匹配到分隔符+换行+书签内容。
  2. 正则模式不匹配:正则试图在同一行内匹配==========+换行符+书签内容,但逐行读取时,换行符不在当前行的内容里,导致(.*)捕获不到任何内容。

解决思路与修正代码:

改成读取整个文件内容,让正则可以跨行匹配。调整正则来匹配分隔符行后紧跟的书签行:

import re

# 读取整个文件内容
with open(r'bookmarks.txt', encoding="utf-8") as f:
    content = f.read()

# 正则匹配:分隔符行 + 换行 + 书签行(非空行)
pattern = re.compile(r'==========\r?\n(.*)', re.MULTILINE)
matches = pattern.findall(content)
count = len(matches)

# 输出结果
for bookmark_line in matches:
    print(f"==========\n{bookmark_line}")
print(f"The amount of notes are: {count}")

说明:

  • re.MULTILINE让正则适配多行文本的匹配逻辑,\r?\n兼容Windows和Unix系统的换行符格式。
  • 用with语句读取文件更安全,会自动关闭文件句柄。
  • 如果需要捕获分隔符+书签行的完整块,也可以把正则调整为r'(==========\r?\n.*)',直接获取包含分隔符的完整内容块。

内容的提问来源于stack exchange,提问作者hector lerma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 17:47:25