Python正则解析日本议会演讲文本的优化需求
日本议会演讲文本解析优化方案
一、先锚定文本结构化特征
日本国会演讲文本有固定格式规律,可先明确核心特征:
- 元数据(届次、日期)集中在文档开头,格式多为「第○○回国会」「平成○○年○月○日/令和○○年○月○日」
- 有效发言的典型格式是「[发言者职位/姓名]:[发言内容]」,而程序性内容(议程、人事变动等)会有明确前缀,比如「議事進行」「人事異動」「採決」等
二、正则匹配优化方案
1. 精准提取元数据
避免泛匹配,针对固定格式写正则:
import re import json def extract_metadata(text): # 匹配议会届次 session_match = re.search(r'第(\d+)回国会', text) # 匹配日期(兼容平成/令和/西历格式) date_match = re.search(r'(平成|令和|\d{4})年(\d{1,2})月(\d{1,2})日', text) return { "议会届次": session_match.group(0) if session_match else None, "日期": date_match.group(0) if date_match else None }
2. 区分发言与程序性内容
先定义程序性内容前缀列表,匹配发言时直接排除这类内容:
def extract_speeches(text): speeches = [] # 可根据测试文件补充更多程序性前缀 procedural_prefixes = ['議事進行', '人事異動', '議案提出', '採決', '閉会', '議案審査'] # 匹配「发言者:内容」格式,跨行内容也能捕获 speech_pattern = re.compile(r'^([^\n:]+):(.+?)(?=\n[^\n:]+:|\Z)', re.MULTILINE | re.DOTALL) for match in speech_pattern.finditer(text): speaker = match.group(1).strip() # 跳过程序性内容 if any(speaker.startswith(prefix) for prefix in procedural_prefixes): continue content = match.group(2).strip().replace('\n', ' ') speeches.append({ "发言者": speaker, "演讲内容": content }) return speeches
3. 整合并导出JSON
def process_text_to_json(text): metadata = extract_metadata(text) speeches = extract_speeches(text) final_result = {**metadata, "演讲内容列表": speeches} with open('parliament_speeches.json', 'w', encoding='utf-8') as f: json.dump(final_result, f, ensure_ascii=False, indent=2)
三、更稳定的状态机解析方案
如果正则仍无法覆盖所有格式变体(比如发言跨行、格式不规范),可以用状态机遍历文本:
def state_machine_parse(text): lines = [line.strip() for line in text.split('\n') if line.strip()] metadata = {"议会届次": None, "日期": None} speeches = [] current_state = "metadata" current_speaker = None current_content = [] procedural_prefixes = ['議事進行', '人事異動', '議案提出', '採決'] for line in lines: if current_state == "metadata": # 提取元数据 session_match = re.search(r'第(\d+)回国会', line) date_match = re.search(r'(平成|令和|\d{4})年(\d{1,2})月(\d{1,2})日', line) if session_match: metadata["议会届次"] = session_match.group(0) if date_match: metadata["日期"] = date_match.group(0) # 切换到发言状态 if ":" in line and not any(line.startswith(p) for p in procedural_prefixes): current_state = "speech" speaker, content = line.split(":", 1) current_speaker = speaker.strip() current_content.append(content.strip()) elif current_state == "speech": # 遇到程序性内容,保存当前发言并切换状态 if any(line.startswith(p) for p in procedural_prefixes): if current_speaker and current_content: speeches.append({ "发言者": current_speaker, "演讲内容": '\n'.join(current_content) }) current_speaker = None current_content = [] current_state = "procedural" # 遇到新发言者,保存上一个发言 elif ":" in line: if current_speaker and current_content: speeches.append({ "发言者": current_speaker, "演讲内容": '\n'.join(current_content) }) speaker, content = line.split(":", 1) current_speaker = speaker.strip() current_content = [content.strip()] # 同一发言者的续行内容 else: current_content.append(line) elif current_state == "procedural": # 回到发言状态 if ":" in line and not any(line.startswith(p) for p in procedural_prefixes): current_state = "speech" speaker, content = line.split(":", 1) current_speaker = speaker.strip() current_content = [content.strip()] # 保存最后一个未完成的发言 if current_speaker and current_content: speeches.append({ "发言者": current_speaker, "演讲内容": '\n'.join(current_content) }) return {**metadata, "演讲内容列表": speeches}
四、调试技巧
- 把测试文件中误识别的程序性内容补充到
procedural_prefixes列表 - 打印正则匹配的所有结果,对比原文本调整正则边界(比如用
^限定行开头,避免跨行误匹配) - 针对漏抓的发言,检查是否是因为发言者与内容的分隔符不是「:」,或者内容跨行未被正则捕获,调整对应逻辑
内容的提问来源于stack exchange,提问作者Ana17
相关产品推荐
相关产品推荐

