PDF解析中提取跨地址与日期的Rel type字段方案问询
问题场景
解析PDF提取Rel type字段时,目标内容Court Order - Own Recognizance被地址(400 Block HANLON WAY)和日期(01/30/2023)拆分换行,导致原正则仅能匹配未换行的完整内容,无法抓取分散的目标值。
示例解析文本
Rel type: Court Order
400 Block HANLON WAY
01/30/2023
- Own Recognizance
当前失效正则代码
Rel type:\s*(.*)
对应页面结构文本
Court Order 400 Block HANLON WAY 01/30/2023 - Own Recognizance
解决方案
1. 先清理干扰内容再合并匹配
先过滤掉地址、日期这类固定格式的干扰项,再合并剩余的Rel type相关内容:
import re # 解析后的原始文本 raw_text = """Rel type: Court Order 400 Block HANLON WAY 01/30/2023 - Own Recognizance""" # 移除地址(数字+Block+街道格式)和日期(MM/DD/YYYY格式) cleaned_text = re.sub(r'\d+ Block [A-Za-z ]+\n|\d{2}/\d{2}/\d{4}\n', '', raw_text) # 合并换行内容并提取目标值 rel_type = re.sub(r'\n', ' ', cleaned_text).replace('Rel type: ', '').strip() print(rel_type) # 输出: Court Order - Own Recognizance
2. 用正则跨换行匹配并跳过干扰项
启用正则的DOTALL模式(让.匹配换行符),直接跳过中间的地址和日期内容:
match = re.search(r'Rel type:\s*(.*?)\s*\d+ Block [A-Za-z ]+\s*\d{2}/\d{2}/\d{4}\s*(.*)', raw_text, re.DOTALL) if match: rel_type = f"{match.group(1).strip()} {match.group(2).strip()}" print(rel_type) # 输出: Court Order - Own Recognizance
3. 基于页面结构直接定位(如果有结构化输出)
如果解析后能拿到页面HTML/XML结构,直接跳过地址、日期标签,合并目标内容:
from bs4 import BeautifulSoup html = """<div class="field"> <label>Rel type:</label> <span>Court Order</span> <span class="address">400 Block HANLON WAY</span> <span class="date">01/30/2023</span> <span>- Own Recognizance</span> </div>""" soup = BeautifulSoup(html, 'html.parser') field_div = soup.find('div', class_='field') # 筛选非地址、非日期的span内容并合并 rel_parts = [span.text.strip() for span in field_div.find_all('span') if not span.get('class') or 'address' not in span.get('class') and 'date' not in span.get('class')] rel_type = ' '.join(rel_parts) print(rel_type) # 输出: Court Order - Own Recognizance
内容的提问来源于stack exchange,提问作者snoozy
相关产品推荐
相关产品推荐

