You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Camelot或同类工具提取PDF表格中的文本与超链接?

提取PDF表格文本与嵌入式超链接的解决方案

Camelot的局限

Camelot的核心设计是解析PDF中的表格结构与文本内容,不支持直接提取嵌入式超链接——因为超链接属于PDF的注解(Annotation)范畴,Camelot并未实现解析这类注解的逻辑。

可行的解决方案

1. Camelot + pdfplumber 组合提取

利用Camelot精准的表格结构提取能力,搭配pdfplumber的PDF注解解析功能,通过坐标匹配将超链接与表格单元格关联:

步骤示例:

  • 用Camelot提取表格文本与单元格位置
import camelot

# 提取第1页的表格
tables = camelot.read_pdf("target.pdf", pages="1")
table_df = tables[0].df
  • 用pdfplumber获取页面中的超链接注解
import pdfplumber

with pdfplumber.open("target.pdf") as pdf:
    page = pdf.pages[0]
    # 过滤出超链接类型的注解
    hyperlinks = [
        annot for annot in page.annots
        if annot.get("subtype") == "Link" and annot.get("uri")
    ]
  • 编写坐标匹配函数,将超链接映射到对应单元格
def match_link_to_cell(cell, hyperlinks):
    # 获取单元格的边界坐标(x1左, y1下, x2右, y2上)
    cell_bbox = (cell.x1, cell.y1, cell.x2, cell.y2)
    for link in hyperlinks:
        link_bbox = link["bbox"]
        # 判断链接是否完全落在单元格范围内(可根据需求调整匹配逻辑)
        if (link_bbox[0] >= cell_bbox[0] and link_bbox[2] <= cell_bbox[2] and
            link_bbox[1] >= cell_bbox[1] and link_bbox[3] <= cell_bbox[3]):
            return link["uri"]
    return None
  • 遍历表格单元格,合并文本与超链接
# 获取表格的单元格对象列表
table_cells = tables[0].cells

for row_idx in range(len(table_df)):
    for col_idx in range(len(table_df.columns)):
        target_cell = table_cells[row_idx][col_idx]
        matched_link = match_link_to_cell(target_cell, hyperlinks)
        if matched_link:
            # 将链接追加到单元格文本后,或自定义格式
            table_df.iloc[row_idx, col_idx] = f"{table_df.iloc[row_idx, col_idx]} | 链接: {matched_link}"

# 输出带链接的表格
print(table_df)

2. 直接使用pdfplumber提取表格与超链接

pdfplumber本身支持表格提取,同时可以直接访问页面注解,无需依赖Camelot:

import pdfplumber

with pdfplumber.open("target.pdf") as pdf:
    page = pdf.pages[0]
    # 提取表格
    tables = page.extract_tables()
    # 获取超链接
    hyperlinks = [
        annot for annot in page.annots
        if annot.get("subtype") == "Link" and annot.get("uri")
    ]

    # 遍历表格单元格匹配链接(逻辑同上述match_link_to_cell函数)
    for table in tables:
        for row_idx, row in enumerate(table):
            for col_idx, cell_text in enumerate(row):
                # 获取单元格边界(需根据pdfplumber的表格单元格坐标调整)
                cell_bbox = page.find_table_cell(row_idx, col_idx)["bbox"]
                # 匹配链接逻辑...

注意事项

  • PDF的坐标系统可能存在y轴方向差异(部分工具以页面底部为原点,部分以顶部为原点),需根据实际情况调整坐标匹配逻辑
  • 若超链接跨多个单元格,或单元格内包含多个超链接,需优化匹配规则以覆盖这类场景
  • 扫描版PDF(图片格式)无法提取超链接,需先进行OCR,但超链接信息会丢失

内容的提问来源于stack exchange,提问作者Amy D

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 16:35:21