如何高效用Python正则分析FOMC会议纪要中的发言归属?
优化FOMC会议纪要的发言者-内容匹配方案
我需要对FOMC会议纪要开展文本分析,识别“谁发言了什么”。已将会议纪要PDF转换为文本,当前计划通过正则提取全大写格式的姓名并拆分发言,代码如下:
names = re.findall("\s\n{2,}[A-Z]{2,}\.*\s*[A-Z]{2,}\.\d*\s",text) speech = re.split("\s\n{2,}[A-Z]{2,}\.*\s*[A-Z]{2,}\.\d*\s",text)随后将这两个列表写入含“names”“speech”两列的CSV,但该方法效率较低,请问是否有更优方案?
附会议纪要样本:
CHAIRMAN BERNANKE. Good afternoon, everybody. PARTICIPANTS. Good afternoon. CHAIRMAN BERNANKE. We need a motion to close our meeting. MR. KOHN. So moved. CHAIRMAN BERNANKE. Thank you. Our meeting today and tomorrow follows the basic sequence we’ve been having recently, but with an important addition, which is that we have a staff presentation on inflation dynamics. We need about two hours for that presentation, I understand, and we’ve thought about it and decided to put it at the end of the meeting so we would have plenty of time to complete our policy decision. But I hope that people will pay attention to the time and make sure we have enough time tomorrow to give appropriate attention to the presentation. In that spirit, why don’t we start directly? Mr. Sack. MR. SACK. Since the last FOMC meeting, financial conditions have generally become more supportive of economic growth.
你的思路方向没问题,但分开用findall和split会做重复的文本遍历,处理长篇会议纪要时效率会打折扣。这里有个更高效的方案,只需要遍历文本一次就能同时提取发言者和对应的内容,还能避免匹配错误:
核心优化点
- 单次遍历完成匹配:用
re.finditer一次性捕获所有发言者和对应的发言内容,避免两次扫描文本。 - 精准覆盖发言者格式:调整正则,适配FOMC常见的发言者头衔(比如
CHAIRMAN XYZ、MR. ABC这类全大写格式)。 - 自动清理杂乱格式:处理跨多行的发言,合并多余换行和空格,直接得到整洁的文本内容。
完整代码示例
import re import csv # 替换成你的会议纪要文本 fomc_text = """CHAIRMAN BERNANKE. Good afternoon, everybody. PARTICIPANTS. Good afternoon. CHAIRMAN BERNANKE. We need a motion to close our meeting. MR. KOHN. So moved. CHAIRMAN BERNANKE. Thank you. Our meeting today and tomorrow follows the basic sequence we’ve been having recently, but with an important addition, which is that we have a staff presentation on inflation dynamics. We need about two hours for that presentation, I understand, and we’ve thought about it and decided to put it at the end of the meeting so we would have plenty of time to complete our policy decision. But I hope that people will pay attention to the time and make sure we have enough time tomorrow to give appropriate attention to the presentation. In that spirit, why don’t we start directly? Mr. Sack. MR. SACK. Since the last FOMC meeting, financial conditions have generally become more supportive of economic growth.""" # 优化后的正则:匹配两个换行后的发言者,捕获到下一个发言者或文本结束 # 适配头衔带点的格式(如MR.)和全头衔格式(如CHAIRMAN) speaker_pattern = re.compile( r'(?<=\n\n)([A-Z]+(?:\.[A-Z]+)?(?:\s[A-Z]+)*\.)\s*(.*?)(?=\n\n[A-Z]+(?:\.[A-Z]+)?(?:\s[A-Z]+)*\.|\Z)', re.DOTALL ) # 一次性提取所有匹配结果 speech_data = [] for match in speaker_pattern.finditer(fomc_text): speaker = match.group(1).strip() # 清理发言里的多余换行、空格,合并成整洁的单行文本 cleaned_speech = re.sub(r'\s+', ' ', match.group(2)).strip() speech_data.append({"names": speaker, "speech": cleaned_speech}) # 写入CSV文件 with open('fomc_speeches.csv', 'w', newline='', encoding='utf-8') as csv_file: writer = csv.DictWriter(csv_file, fieldnames=["names", "speech"]) writer.writeheader() writer.writerows(speech_data)
为什么这个方案更好?
- 效率更高:只遍历文本一次,相比你原来的两次遍历(
findall+split),处理大文件时速度提升明显。 - 匹配更准确:正则里的正向预查
(?=\n\n[A-Z]+...|\Z)会精准定位发言的结束位置,不会把发言中间的内容错误拆分;同时适配了带点的头衔格式,避免漏匹配。 - 更少的后续处理:自动清理了跨多行的换行和多余空格,直接得到可以用的整洁文本,不用再额外写处理逻辑。
如果你的会议纪要里还有带数字的特殊格式发言者,只需在正则的发言者部分加上\d*即可微调适配。
内容的提问来源于stack exchange,提问作者user9767961
相关产品推荐
相关产品推荐

