如何用R从文本中提取指定信息并解决重复匹配问题
批量处理法律文本:提取身份信息+婚姻法规检测方案
需求拆解
- 从每条文本中提取首次出现的原告(plaintiff)、被告(defendant)身份语句,要求语句包含出生日期(born)和住址(lives)信息,排除无关表述
- 检测文本是否包含《marriage law》,存在标记为
T,不存在标记为F - 输出结构化结果表
实现代码(Python)
用正则表达式精准匹配目标内容,结合pandas批量处理生成结果表:
import re import pandas as pd def process_legal_text(text): # 提取原告首次有效身份信息 plaintiff_pattern = r"The plaintiff.*?born.*?lives.*?\." plaintiff_match = re.search(plaintiff_pattern, text, re.DOTALL) plaintiff_info = plaintiff_match.group().strip() if plaintiff_match else None # 提取被告首次有效身份信息 defendant_pattern = r"The defendant.*?born.*?lives.*?\." defendant_match = re.search(defendant_pattern, text, re.DOTALL) defendant_info = defendant_match.group().strip() if defendant_match else None # 检测《marriage law》 has_marriage_law = "T" if "《marriage law》" in text else "F" return { "plaintiff_info": plaintiff_info, "defendant_info": defendant_info, "has_marriage_law": has_marriage_law } # 示例文本测试 sample_text = """Ganluo County People's Court of X Province。The plaintiff X, female, born on May, 1980, lives in X County, X Province。The defendant X, male, born on May, 1971, lives in X County, X Province。 It is a divorce dispute, according to 《marriage law》on June 21, 2016。""" # 批量处理(假设3000条文本存在text_list列表中) # text_list = [text1, text2, ..., text3000] # results = [process_legal_text(text) for text in text_list] # 生成结果表 result_df = pd.DataFrame([process_legal_text(sample_text)]) print(result_df)
代码说明
- 正则匹配逻辑:
- 用
.*?非贪婪匹配确保只捕获从"The plaintiff/defendant"到第一个包含born和lives的句子结尾(.) re.DOTALL参数允许.匹配换行符,避免跨段落的文本被截断
- 用
- 批量处理:将所有文本存入列表后,通过列表推导式批量调用处理函数
- 结果输出:用pandas将结果转为DataFrame,方便导出为Excel/CSV格式
测试结果(示例文本)
| plaintiff_info | defendant_info | has_marriage_law |
|---|---|---|
| The plaintiff X, female, born on May, 1980, lives in X County, X Province。 | The defendant X, male, born on May, 1971, lives in X County, X Province。 | T |
内容的提问来源于stack exchange,提问作者Xinyan LIU
相关产品推荐
相关产品推荐

