适配三类PDF的通用表格数据提取正则表达式需求
通用正则适配三类PDF表格数据提取
问题分析
三类PDF的表格行存在以下共性与差异:
- 名称列:可能以数字开头(如
1 XYZ Corp、15/92 Advertisement Inc.),也可能无前置数字(如FC LC-New York);名称包含字母、空格、斜杠、连字符、点等特殊字符,大小写混合。 - 分配额列:格式覆盖纯小数(
6.00)、带$符号($2.0)、带k/m后缀(500.0k、2.0m)三种情况。
通用正则表达式
针对上述场景,设计通用正则匹配每行的名称与分配额:
import re # 匹配名称(账户/贷方)和分配额的通用正则 row_regex = re.compile(r'^\s*(?:\S+\s+)?(.+?)\s+(\$?\d+(?:\.\d+)?(?:[km])?)\s*$')
正则拆解
^\s*:匹配行首任意空白字符(?:\S+\s+)?:可选的前置非空白内容(如数字1、15/92),不捕获该部分(.+?):非贪婪匹配名称内容(账户/贷方名称),捕获为第1组\s+:匹配名称与分配额之间的空白分隔(\$?\d+(?:\.\d+)?(?:[km])?):匹配分配额,捕获为第2组:\$?:可选的$符号\d+(?:\.\d+)?:整数或小数部分(?:[km])?:可选的k/m后缀
\s*$:匹配行尾任意空白字符
结合现有代码的使用示例
利用你已有的表头识别逻辑,遍历表格行提取数据:
import re import json row_regex = re.compile(r'^\s*(?:\S+\s+)?(.+?)\s+(\$?\d+(?:\.\d+)?(?:[km])?)\s*$') # 匹配三类表头的正则 header_regex = re.compile(r'(Accounts|Account|Lender)\s+Allocation') # 替换为你从PDF提取的行数据 lines = [ "| Accounts | Allocation |", "|----------|------------|", "| 1 XYZ Corp | 6.00 |", "| 2 BCF | 3.00 |", "| 3 Barings | 2.50 |", "", "| Account | Allocation |", "|---------|------------|", "| 1 Amep | $2.0 |", "| 2 Asset Pioneer | $13.0 |", "| 3 Creed Partners | $35.5 |", "", "| Lender | Allocation |", "|--------|------------|", "| 15/92 Advertisement Inc. | 500.0k |", "| FC LC-New York | 2.0m |", "| ABE PARTNERS INC | 5.0m |" ] table_data = [] current_table = [] table_start = None current_header = None for i, line in enumerate(lines): # 识别表头并记录当前列类型 header_match = header_regex.search(line) if header_match: table_start = i + 2 # 跳过表头分隔线行 current_header = header_match.group(1) current_table = [] # 处理表格行 elif table_start is not None and i >= table_start: line_clean = line.strip().strip('|').strip() if not line_clean: # 表格结束,存入总数据 table_data.append({ "header": current_header, "rows": current_table }) table_start = None current_header = None continue # 用通用正则匹配行数据 row_match = row_regex.match(line_clean) if row_match: name = row_match.group(1).strip() allocation = row_match.group(2).strip() current_table.append({ current_header.lower(): name, "allocation": allocation }) # 输出JSON格式结果 print(json.dumps(table_data, indent=2))
适配扩展说明
- 若分配额后缀存在大小写差异(如
K/M),可将正则中的[km]改为[kKmM]; - 若PDF提取的行存在单元格换行,需先对行进行合并预处理;
- 若名称中包含数字(示例未覆盖),可通过表头列的固定位置拆分数据,避免误匹配。
内容的提问来源于stack exchange,提问作者Jess
相关产品推荐
相关产品推荐

