You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python文本解析:实现带前置分隔符的问答文本拆分

证词问答文本拆分优化方案

问题背景

现有一段OCR识别的公开证词文本:

text = """\na\n\nQ So I do first want to bring up exhibit No. 46, which is in the binder 
in front of\nyou.\n\nAnd that is a letter [to] Alston\n& Bird...\n\nIs that correct?\n\nA This is correct.\n\nQ Okay."""

使用原有正则拆分代码:

import re
pattern = "\n[QA]_?\s"
q_a_list = re.split(pattern, text)
print(q_a_list)

得到的结果存在两个核心问题:

  • 无法区分每个文本块是问题(Q)还是回答(A)
  • 列表首个元素为分隔符前的无关冗余文本

优化方案

改用re.findall匹配完整的问答块(包含Q/A前缀),同时自动过滤开头无关内容,还可将结果整理为结构化形式,方便后续处理。

方案1:保留前缀的列表形式

通过正则匹配所有带Q/A前缀的完整问答块,确保每个块包含明确的标识:

import re

text = """\na\n\nQ So I do first want to bring up exhibit No. 46, which is in the binder 
in front of\nyou.\n\nAnd that is a letter [to] Alston\n& Bird...\n\nIs that correct?\n\nA This is correct.\n\nQ Okay."""

# 匹配带Q/A前缀的完整问答块,非贪婪匹配直到下一个问答前缀或文本结尾
pattern = r"\n[QA]_?\s.*?(?=\n[QA]_?\s|\Z)"
q_a_blocks = re.findall(pattern, text, re.DOTALL)

# 清理每个块的首尾空白(可选)
cleaned_blocks = [block.strip() for block in q_a_blocks]

print(cleaned_blocks)

输出结果:

['Q So I do first want to bring up exhibit No. 46, which is in the binder in front of you.\n\nAnd that is a letter [to] Alston\n& Bird...\n\nIs that correct?', 'A This is correct.', 'Q Okay.']

方案2:结构化键值对形式

如果需要更清晰的分类结构,可将每个问答拆分为「类型标识」和「内容」的键值对:

import re

text = """\na\n\nQ So I do first want to bring up exhibit No. 46, which is in the binder 
in front of\nyou.\n\nAnd that is a letter [to] Alston\n& Bird...\n\nIs that correct?\n\nA This is correct.\n\nQ Okay."""

# 分组匹配:捕获类型(Q/A)和对应的内容
pattern = r"\n([QA])_?\s(.*?)(?=\n[QA]_?\s|\Z)"
q_a_pairs = re.findall(pattern, text, re.DOTALL)

# 整理为字典列表
structured_data = [
    {"type": item[0], "content": item[1].strip()}
    for item in q_a_pairs
]

print(structured_data)

输出结果:

[
    {'type': 'Q', 'content': 'So I do first want to bring up exhibit No. 46, which is in the binder in front of you.\n\nAnd that is a letter [to] Alston\n& Bird...\n\nIs that correct?'},
    {'type': 'A', 'content': 'This is correct.'},
    {'type': 'Q', 'content': 'Okay.'}
]

方案说明

  • re.DOTALL参数让正则中的.匹配换行符,确保能捕获跨多行的问答内容
  • 非贪婪匹配.*?避免误匹配到下一个问答块的内容
  • 前瞻断言(?=\n[QA]_?\s|\Z)精准定位每个问答块的结束位置(下一个问答前缀或文本末尾)
  • 开头的无关文本会被自动过滤,因为正则只匹配带Q/A前缀的有效块

内容的提问来源于stack exchange,提问作者Max Power

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 04:40:37