You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则实现多行文本中发言人与内容的精准拆分

解决Python中用re捕获多行发言人及内容的问题

问题场景

在Python 3.9.X环境下,需要处理一段包含多位发言人的多行字符串:

sample_string = """STEVE SMITH, AMERICAN DAD : Good morning, good 
afternoon, usa . Before I hand over to  Homer, I want to give a quick reminder of the cartoons we are making.  

Numbers in the presentation today. Our focus is now on reported num bers, but we will call out and 
specify notable items . we like films and ultimately benefit your schedule going 
forward . Homer, over to you.  

HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD:  Thanks, Steve, and good morning in China, 
good afternoon in welcome to our viewership results call . Beans is 
going to lead the presentation, but I’d like to make some opening comments.  

We’ve announced about 1000 hours viewing time, so our strategy is working."""

发言人以全大写名称+冒号标识,发言内容跨多行。需要捕获两组内容:发言人信息、对应发言文本,最终输出为元组列表,示例如下:

output = [
    ("STEVE SMITH, AMERICAN DAD", "Good morning, good \nafternoon, usa . Before I hand over to  Homer, I want to give a quick reminder of the cartoons we are making.  \n\nNumbers in the presentation today. Our focus is now on reported num bers, but we will call out and \nspecify notable items . we like films and ultimately benefit your schedule going \nforward . Homer, over to you.  "),
    ("HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD", "Thanks, Steve, and good morning in China, \ngood afternoon in welcome to our viewership results call . Beans is \ngoing to lead the presentation, but I’d like to make some opening comments.  \n\nWe’ve announced about 1000 hours viewing time, so our strategy is working.")
]

原正则表达式^([^a-z:]+?)\s*:\s*(.|\n?)无法在遇到新发言人时停止捕获,需要优化。

解决方案

使用正向预查限定发言内容的结束边界,结合re.DOTALL和re.MULTILINE标志实现精准匹配:

正则表达式

^([A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z]|$)

Python代码实现

import re

sample_string = """STEVE SMITH, AMERICAN DAD : Good morning, good 
afternoon, usa . Before I hand over to  Homer, I want to give a quick reminder of the cartoons we are making.  

Numbers in the presentation today. Our focus is now on reported num bers, but we will call out and 
specify notable items . we like films and ultimately benefit your schedule going 
forward . Homer, over to you.  

HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD:  Thanks, Steve, and good morning in China, 
good afternoon in welcome to our viewership results call . Beans is 
going to lead the presentation, but I’d like to make some opening comments.  

We’ve announced about 1000 hours viewing time, so our strategy is working."""

# 编译正则,启用多行模式和DOTALL模式
pattern = re.compile(r'^([A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z]|$)', re.DOTALL | re.MULTILINE)
# 匹配所有结果
matches = pattern.findall(sample_string)

# 输出结果
for speaker, content in matches:
    print(f"发言人: {speaker.strip()}")
    print(f"内容:\n{content.strip()}\n---")

正则规则解释

  • ^([A-Z\s,]+?):匹配行首的全大写发言人信息(允许空格、逗号),+?非贪婪匹配避免过度捕获无关内容
  • \s*:\s*:匹配冒号前后的任意空白字符,兼容冒号前后空格不一致的情况
  • (.*?):非贪婪匹配发言内容,re.DOTALL让.能够匹配换行符,支持跨多行内容捕获
  • (?=\n[A-Z]|$):正向预查边界,当检测到换行后接大写字母(新发言人开头)或字符串结尾时,停止当前发言内容的捕获

内容的提问来源于stack exchange,提问作者Beans On Toast

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 07:48:17