Python正则实现多行文本中发言人与内容的精准拆分
解决Python中用re捕获多行发言人及内容的问题
问题场景
在Python 3.9.X环境下,需要处理一段包含多位发言人的多行字符串:
sample_string = """STEVE SMITH, AMERICAN DAD : Good morning, good afternoon, usa . Before I hand over to Homer, I want to give a quick reminder of the cartoons we are making. Numbers in the presentation today. Our focus is now on reported num bers, but we will call out and specify notable items . we like films and ultimately benefit your schedule going forward . Homer, over to you. HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD: Thanks, Steve, and good morning in China, good afternoon in welcome to our viewership results call . Beans is going to lead the presentation, but I’d like to make some opening comments. We’ve announced about 1000 hours viewing time, so our strategy is working."""
发言人以全大写名称+冒号标识,发言内容跨多行。需要捕获两组内容:发言人信息、对应发言文本,最终输出为元组列表,示例如下:
output = [ ("STEVE SMITH, AMERICAN DAD", "Good morning, good \nafternoon, usa . Before I hand over to Homer, I want to give a quick reminder of the cartoons we are making. \n\nNumbers in the presentation today. Our focus is now on reported num bers, but we will call out and \nspecify notable items . we like films and ultimately benefit your schedule going \nforward . Homer, over to you. "), ("HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD", "Thanks, Steve, and good morning in China, \ngood afternoon in welcome to our viewership results call . Beans is \ngoing to lead the presentation, but I’d like to make some opening comments. \n\nWe’ve announced about 1000 hours viewing time, so our strategy is working.") ]
原正则表达式^([^a-z:]+?)\s*:\s*(.|\n?)无法在遇到新发言人时停止捕获,需要优化。
解决方案
使用正向预查限定发言内容的结束边界,结合re.DOTALL和re.MULTILINE标志实现精准匹配:
正则表达式
^([A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z]|$)
Python代码实现
import re sample_string = """STEVE SMITH, AMERICAN DAD : Good morning, good afternoon, usa . Before I hand over to Homer, I want to give a quick reminder of the cartoons we are making. Numbers in the presentation today. Our focus is now on reported num bers, but we will call out and specify notable items . we like films and ultimately benefit your schedule going forward . Homer, over to you. HOMER SIMPSON, HEAD OF SIMPSON HOUSEHOLD: Thanks, Steve, and good morning in China, good afternoon in welcome to our viewership results call . Beans is going to lead the presentation, but I’d like to make some opening comments. We’ve announced about 1000 hours viewing time, so our strategy is working.""" # 编译正则,启用多行模式和DOTALL模式 pattern = re.compile(r'^([A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z]|$)', re.DOTALL | re.MULTILINE) # 匹配所有结果 matches = pattern.findall(sample_string) # 输出结果 for speaker, content in matches: print(f"发言人: {speaker.strip()}") print(f"内容:\n{content.strip()}\n---")
正则规则解释
^([A-Z\s,]+?):匹配行首的全大写发言人信息(允许空格、逗号),+?非贪婪匹配避免过度捕获无关内容\s*:\s*:匹配冒号前后的任意空白字符,兼容冒号前后空格不一致的情况(.*?):非贪婪匹配发言内容,re.DOTALL让.能够匹配换行符,支持跨多行内容捕获(?=\n[A-Z]|$):正向预查边界,当检测到换行后接大写字母(新发言人开头)或字符串结尾时,停止当前发言内容的捕获
内容的提问来源于stack exchange,提问作者Beans On Toast
相关产品推荐
相关产品推荐

