You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

适配三类PDF的通用表格数据提取正则表达式需求

通用正则适配三类PDF表格数据提取

问题分析

三类PDF的表格行存在以下共性与差异:

  • 名称列:可能以数字开头(如1 XYZ Corp、15/92 Advertisement Inc.),也可能无前置数字(如FC LC-New York);名称包含字母、空格、斜杠、连字符、点等特殊字符,大小写混合。
  • 分配额列:格式覆盖纯小数(6.00)、带$符号($2.0)、带k/m后缀(500.0k、2.0m)三种情况。

通用正则表达式

针对上述场景,设计通用正则匹配每行的名称与分配额:

import re

# 匹配名称(账户/贷方)和分配额的通用正则
row_regex = re.compile(r'^\s*(?:\S+\s+)?(.+?)\s+(\$?\d+(?:\.\d+)?(?:[km])?)\s*$')

正则拆解

  • ^\s*:匹配行首任意空白字符
  • (?:\S+\s+)?:可选的前置非空白内容(如数字1、15/92),不捕获该部分
  • (.+?):非贪婪匹配名称内容(账户/贷方名称),捕获为第1组
  • \s+:匹配名称与分配额之间的空白分隔
  • (\$?\d+(?:\.\d+)?(?:[km])?):匹配分配额,捕获为第2组:
    • \$?:可选的$符号
    • \d+(?:\.\d+)?:整数或小数部分
    • (?:[km])?:可选的k/m后缀
  • \s*$:匹配行尾任意空白字符

结合现有代码的使用示例

利用你已有的表头识别逻辑,遍历表格行提取数据:

import re
import json

row_regex = re.compile(r'^\s*(?:\S+\s+)?(.+?)\s+(\$?\d+(?:\.\d+)?(?:[km])?)\s*$')
# 匹配三类表头的正则
header_regex = re.compile(r'(Accounts|Account|Lender)\s+Allocation')

# 替换为你从PDF提取的行数据
lines = [
    "| Accounts | Allocation |",
    "|----------|------------|",
    "| 1 XYZ Corp | 6.00 |",
    "| 2 BCF | 3.00 |",
    "| 3 Barings | 2.50 |",
    "",
    "| Account | Allocation |",
    "|---------|------------|",
    "| 1 Amep | $2.0 |",
    "| 2 Asset Pioneer | $13.0 |",
    "| 3 Creed Partners | $35.5 |",
    "",
    "| Lender | Allocation |",
    "|--------|------------|",
    "| 15/92 Advertisement Inc. | 500.0k |",
    "| FC LC-New York | 2.0m |",
    "| ABE PARTNERS INC | 5.0m |"
]

table_data = []
current_table = []
table_start = None
current_header = None

for i, line in enumerate(lines):
    # 识别表头并记录当前列类型
    header_match = header_regex.search(line)
    if header_match:
        table_start = i + 2  # 跳过表头分隔线行
        current_header = header_match.group(1)
        current_table = []
    # 处理表格行
    elif table_start is not None and i >= table_start:
        line_clean = line.strip().strip('|').strip()
        if not line_clean:
            # 表格结束,存入总数据
            table_data.append({
                "header": current_header,
                "rows": current_table
            })
            table_start = None
            current_header = None
            continue
        # 用通用正则匹配行数据
        row_match = row_regex.match(line_clean)
        if row_match:
            name = row_match.group(1).strip()
            allocation = row_match.group(2).strip()
            current_table.append({
                current_header.lower(): name,
                "allocation": allocation
            })

# 输出JSON格式结果
print(json.dumps(table_data, indent=2))

适配扩展说明

  • 若分配额后缀存在大小写差异(如K/M),可将正则中的[km]改为[kKmM];
  • 若PDF提取的行存在单元格换行,需先对行进行合并预处理;
  • 若名称中包含数字(示例未覆盖),可通过表头列的固定位置拆分数据,避免误匹配。

内容的提问来源于stack exchange,提问作者Jess

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:30:37