如何用Python正则表达式拆分发言文本中的发言人和内容?
正则拆分发言人与内容方案
正则表达式模式
使用以下正则(需开启多行模式(MULTILINE)和点匹配换行模式(DOTALL)):
^([^:]+?)\s*:\s*(.*?)(?=\n\n[^:]+?:|\Z)
正则各部分说明
^:匹配每行的起始位置(多行模式下生效)([^:]+?):捕获发言人信息,匹配到第一个冒号前的所有内容(非贪婪模式避免过度匹配)\s*:\s*:匹配冒号及前后任意数量的空白字符(适配原文本中冒号前后的空格差异)(.*?):捕获发言内容,非贪婪模式匹配到下一个分隔符前的所有内容(?=\n\n[^:]+?:|\Z):正向预查,触发匹配结束的条件是「两个换行+新的发言人标识」或「文本末尾」
实际使用示例(Python)
import re text = """JOHN SMITH, GLOBAL HEAD OF YOUTUBE : Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show. BOBBY DUDE, GROUP FROM FACEBOOK: Thanks, john smith lets talk about movies and films we watch when we are bored parents. """ pattern = re.compile(r'^([^:]+?)\s*:\s*(.*?)(?=\n\n[^:]+?:|\Z)', re.M | re.S) matches = pattern.findall(text) # 输出拆分结果并翻译 for idx, (speaker, content) in enumerate(matches, 1): # 翻译发言人信息 speaker_cn = speaker.replace("JOHN SMITH, GLOBAL HEAD OF YOUTUBE", "约翰·史密斯,YouTube全球负责人")\ .replace("BOBBY DUDE, GROUP FROM FACEBOOK", "鲍比·杜德,Facebook团队成员") # 翻译发言内容 content_cn = content.replace("Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show.", # 正则表达式拆分发言人和发言内容方案 ### 适用的正则模式 针对你提供的文本格式,可使用以下正则表达式(启用多行模式): ```regex (?m)^([A-Z][A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z][A-Z\s,]+:|\Z)
模式逐段解释
(?m):启用多行模式,让^匹配每一行的开头,确保只从行首识别发言人^([A-Z][A-Z\s,]+?):匹配发言人信息,从行首大写字母开始,匹配大写字母、空格、逗号,用非贪婪模式避免过度匹配到无关内容\s*::匹配冒号前的任意数量空格(包括0个,兼容不同格式的冒号写法)\s*:匹配冒号后的任意数量空格,清理发言内容前的冗余空格(.*?):非贪婪匹配发言内容,避免把后续发言人信息包含进来(?=\n[A-Z][A-Z\s,]+:|\Z):正向预查,确保内容截止到下一个行首的发言人标记(大写字母开头+冒号)或文本结束位置
代码示例(Python)
import re text = """JOHN SMITH, GLOBAL HEAD OF YOUTUBE : Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show. BOBBY DUDE, GROUP FROM FACEBOOK: Thanks, john smith lets talk about movies and films we watch when we are bored parents. """ pattern = r'(?m)^([A-Z][A-Z\s,]+?)\s*:\s*(.*?)(?=\n[A-Z][A-Z\s,]+:|\Z)' matches = re.findall(pattern, text, re.DOTALL) for idx, (speaker, content) in enumerate(matches, 1): # 清理内容中的多余空行和首尾空格 cleaned_content = re.sub(r'\n\s*\n', '\n', content).strip() print(f'string{idx}: (speaker = {speaker.strip()}, text = {cleaned_content} )')
输出结果
string1: (speaker = JOHN SMITH, GLOBAL HEAD OF YOUTUBE, text = Good morning, good afternoon, everyone . Before I hand over to facebook, I want to give a quick reminder of the reporting changes that have taken effect this filming of a tv show. ) string2: (speaker = BOBBY DUDE, GROUP FROM FACEBOOK, text = Thanks, john smith lets talk about movies and films we watch when we are bored parents. )
内容的提问来源于stack exchange,提问作者DreadPirateRoberts
相关产品推荐
相关产品推荐

