You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式拆分发言文本中的发言人和内容?

正则拆分发言人与内容方案

正则表达式模式

使用以下正则(需开启多行模式(MULTILINE)和点匹配换行模式(DOTALL)):

^([^:]+?)\s*:\s*(.*?)(?=\n\n[^:]+?:|\Z)

正则各部分说明

  • ^:匹配每行的起始位置(多行模式下生效)
  • ([^:]+?):捕获发言人信息,匹配到第一个冒号前的所有内容(非贪婪模式避免过度匹配)
  • \s*:\s*:匹配冒号及前后任意数量的空白字符(适配原文本中冒号前后的空格差异)
  • (.*?):捕获发言内容,非贪婪模式匹配到下一个分隔符前的所有内容
  • (?=\n\n[^:]+?:|\Z):正向预查,触发匹配结束的条件是「两个换行+新的发言人标识」或「文本末尾」

实际使用示例(Python)

import re

text = """JOHN SMITH, GLOBAL HEAD OF YOUTUBE : Good morning, good 
afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting 
changes that have taken effect this filming of a tv show.  
 
 
BOBBY DUDE, GROUP FROM FACEBOOK:     Thanks, john smith lets talk about movies and films we watch when we are bored parents. 
"""

pattern = re.compile(r'^([^:]+?)\s*:\s*(.*?)(?=\n\n[^:]+?:|\Z)', re.M | re.S)
matches = pattern.findall(text)

# 输出拆分结果并翻译
for idx, (speaker, content) in enumerate(matches, 1):
    # 翻译发言人信息
    speaker_cn = speaker.replace("JOHN SMITH, GLOBAL HEAD OF YOUTUBE", "约翰·史密斯,YouTube全球负责人")\
                        .replace("BOBBY DUDE, GROUP FROM FACEBOOK", "鲍比·杜德,Facebook团队成员")
    # 翻译发言内容
    content_cn = content.replace("Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show.",
# 正则表达式拆分发言人和发言内容方案

### 适用的正则模式
针对你提供的文本格式,可使用以下正则表达式(启用多行模式):
```regex
(?m)^([A-Z][A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z][A-Z\s,]+:|\Z)

模式逐段解释

  • (?m):启用多行模式,让^匹配每一行的开头,确保只从行首识别发言人
  • ^([A-Z][A-Z\s,]+?):匹配发言人信息,从行首大写字母开始,匹配大写字母、空格、逗号,用非贪婪模式避免过度匹配到无关内容
  • \s*::匹配冒号前的任意数量空格(包括0个,兼容不同格式的冒号写法)
  • \s*:匹配冒号后的任意数量空格,清理发言内容前的冗余空格
  • (.*?):非贪婪匹配发言内容,避免把后续发言人信息包含进来
  • (?=\n[A-Z][A-Z\s,]+:|\Z):正向预查,确保内容截止到下一个行首的发言人标记(大写字母开头+冒号)或文本结束位置

代码示例(Python)

import re

text = """JOHN SMITH, GLOBAL HEAD OF YOUTUBE : Good morning, good 
afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting 
changes that have taken effect this filming of a tv show.  
 
 
BOBBY DUDE, GROUP FROM FACEBOOK:     Thanks, john smith lets talk about movies and films we watch when we are bored parents. 
"""

pattern = r'(?m)^([A-Z][A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z][A-Z\s,]+:|\Z)'
matches = re.findall(pattern, text, re.DOTALL)

for idx, (speaker, content) in enumerate(matches, 1):
    # 清理内容中的多余空行和首尾空格
    cleaned_content = re.sub(r'\n\s*\n', '\n', content).strip()
    print(f'string{idx}: (speaker = {speaker.strip()}, text = {cleaned_content} )')

输出结果

string1: (speaker = JOHN SMITH, GLOBAL HEAD OF YOUTUBE, text = Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show. )
string2: (speaker = BOBBY DUDE, GROUP FROM FACEBOOK, text = Thanks, john smith lets talk about movies and films we watch when we are bored parents. )

内容的提问来源于stack exchange,提问作者DreadPirateRoberts

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 06:35:13