You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用re模块匹配以Q:/A:开头、截止下一个Q:/A:的跨多行字符串

Python正则提取跨多行问答块解决方案

原代码问题原因

  • 你使用的[\w\W]*是贪婪匹配规则,会尽可能匹配最长的字符串,因此会直接从第一个匹配到的Q: /A: 一直吞到整个文本的末尾,自然只能得到一整个匹配结果
  • 你把下一个问答的开头标识(A: |Q: |$)写在了普通捕获组中,这部分会被当作匹配结果的一部分被消费,不仅会把下一个问答的开头包含进当前结果,还会导致下一轮匹配跳过该开头标识

正确实现代码

import re

string = "Q: This is a question. \nQ: This is a 2nd question \non two lines. \n\nA: This is an answer. \nA: This is a 2nd answer \non two lines.\nQ: Here's another question. \nA: And another answer."

# 正则说明:
# [QA]:  匹配Q: 或A: 开头的标识
# .*? 非贪婪匹配任意字符,配合re.DOTALL可匹配换行符
# (?=\s*[QA]: |\Z) 正向零宽预查,匹配到下一个问答开头或文本末尾时停止,不消费后续字符
pattern = re.compile(r'[QA]: .*?(?=\s*[QA]: |\Z)', re.DOTALL)

matches = pattern.finditer(string)
for match in matches:
    print('-', match.group(0).strip()) # strip()可去掉首尾多余空行,不需要可以删除

运行输出效果

- Q: This is a question.
- Q: This is a 2nd question 
on two lines.
- A: This is an answer.
- A: This is a 2nd answer 
on two lines.
- Q: Here's another question.
- A: And another answer.

扩展说明

如果需要区分提取问题和答案,可以给开头加捕获组:
re.compile(r'([QA]): (.*?)(?=\s*[QA]: |\Z)', re.DOTALL)
遍历的时候可以通过match.group(1)获取是Q还是A,match.group(2)获取内容本身,更方便后续存入问答对结构。

内容的提问来源于stack exchange,提问作者Mark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 02:36:03