You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则解析日本议会演讲文本的优化需求

日本议会演讲文本解析优化方案

一、先锚定文本结构化特征

日本国会演讲文本有固定格式规律,可先明确核心特征:

  • 元数据(届次、日期)集中在文档开头,格式多为「第○○回国会」「平成○○年○月○日/令和○○年○月○日」
  • 有效发言的典型格式是「[发言者职位/姓名]:[发言内容]」,而程序性内容(议程、人事变动等)会有明确前缀,比如「議事進行」「人事異動」「採決」等

二、正则匹配优化方案

1. 精准提取元数据

避免泛匹配,针对固定格式写正则:

import re
import json

def extract_metadata(text):
    # 匹配议会届次
    session_match = re.search(r'第(\d+)回国会', text)
    # 匹配日期(兼容平成/令和/西历格式)
    date_match = re.search(r'(平成|令和|\d{4})年(\d{1,2})月(\d{1,2})日', text)
    return {
        "议会届次": session_match.group(0) if session_match else None,
        "日期": date_match.group(0) if date_match else None
    }

2. 区分发言与程序性内容

先定义程序性内容前缀列表,匹配发言时直接排除这类内容:

def extract_speeches(text):
    speeches = []
    # 可根据测试文件补充更多程序性前缀
    procedural_prefixes = ['議事進行', '人事異動', '議案提出', '採決', '閉会', '議案審査']
    # 匹配「发言者:内容」格式,跨行内容也能捕获
    speech_pattern = re.compile(r'^([^\n:]+):(.+?)(?=\n[^\n:]+:|\Z)', re.MULTILINE | re.DOTALL)
    
    for match in speech_pattern.finditer(text):
        speaker = match.group(1).strip()
        # 跳过程序性内容
        if any(speaker.startswith(prefix) for prefix in procedural_prefixes):
            continue
        content = match.group(2).strip().replace('\n', ' ')
        speeches.append({
            "发言者": speaker,
            "演讲内容": content
        })
    return speeches

3. 整合并导出JSON

def process_text_to_json(text):
    metadata = extract_metadata(text)
    speeches = extract_speeches(text)
    final_result = {**metadata, "演讲内容列表": speeches}
    
    with open('parliament_speeches.json', 'w', encoding='utf-8') as f:
        json.dump(final_result, f, ensure_ascii=False, indent=2)

三、更稳定的状态机解析方案

如果正则仍无法覆盖所有格式变体(比如发言跨行、格式不规范),可以用状态机遍历文本:

def state_machine_parse(text):
    lines = [line.strip() for line in text.split('\n') if line.strip()]
    metadata = {"议会届次": None, "日期": None}
    speeches = []
    current_state = "metadata"
    current_speaker = None
    current_content = []
    procedural_prefixes = ['議事進行', '人事異動', '議案提出', '採決']

    for line in lines:
        if current_state == "metadata":
            # 提取元数据
            session_match = re.search(r'第(\d+)回国会', line)
            date_match = re.search(r'(平成|令和|\d{4})年(\d{1,2})月(\d{1,2})日', line)
            if session_match:
                metadata["议会届次"] = session_match.group(0)
            if date_match:
                metadata["日期"] = date_match.group(0)
            # 切换到发言状态
            if ":" in line and not any(line.startswith(p) for p in procedural_prefixes):
                current_state = "speech"
                speaker, content = line.split(":", 1)
                current_speaker = speaker.strip()
                current_content.append(content.strip())
        
        elif current_state == "speech":
            # 遇到程序性内容,保存当前发言并切换状态
            if any(line.startswith(p) for p in procedural_prefixes):
                if current_speaker and current_content:
                    speeches.append({
                        "发言者": current_speaker,
                        "演讲内容": '\n'.join(current_content)
                    })
                    current_speaker = None
                    current_content = []
                current_state = "procedural"
            # 遇到新发言者,保存上一个发言
            elif ":" in line:
                if current_speaker and current_content:
                    speeches.append({
                        "发言者": current_speaker,
                        "演讲内容": '\n'.join(current_content)
                    })
                speaker, content = line.split(":", 1)
                current_speaker = speaker.strip()
                current_content = [content.strip()]
            # 同一发言者的续行内容
            else:
                current_content.append(line)
        
        elif current_state == "procedural":
            # 回到发言状态
            if ":" in line and not any(line.startswith(p) for p in procedural_prefixes):
                current_state = "speech"
                speaker, content = line.split(":", 1)
                current_speaker = speaker.strip()
                current_content = [content.strip()]
    
    # 保存最后一个未完成的发言
    if current_speaker and current_content:
        speeches.append({
            "发言者": current_speaker,
            "演讲内容": '\n'.join(current_content)
        })
    
    return {**metadata, "演讲内容列表": speeches}

四、调试技巧

  1. 把测试文件中误识别的程序性内容补充到procedural_prefixes列表
  2. 打印正则匹配的所有结果,对比原文本调整正则边界(比如用^限定行开头,避免跨行误匹配)
  3. 针对漏抓的发言,检查是否是因为发言者与内容的分隔符不是「:」,或者内容跨行未被正则捕获,调整对应逻辑

内容的提问来源于stack exchange,提问作者Ana17

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 19:23:19