You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyPDF2/PyPDF4提取PDF中URL关联的文本内容?

问题描述

我有一份包含超链接的PDF文档,需要提取URL对应的关联文本。使用PyPDF2和PyPDF4库时,仅能成功提取URL,无法获取绑定的文本内容——比如PDF里有一段绑定了URL的文本“Check this link out”,只能提取到URL,却拿不到这段关联文本。

当前使用的PyPDF4代码如下:

import PyPDF4
import requests

# Open the PDF file
pdf_file = open('abc.pdf', 'rb')

# Create a PDF reader object
pdf_reader = PyPDF4.PdfFileReader(pdf_file)

# Loop through each page of the PDF
for page_num in range(len(pdf_reader.pages)):
    # Get the page object
    page = pdf_reader.pages[page_num]

    # Extract the annotations from the page
    annotations = page.get('/Annots')

    # If there are no annotations, skip to the next page
    if not annotations:
        continue

    # Loop through each annotation
    for annotation in annotations:
        # Get the annotation dictionary
        annotation_dict = annotation.getObject()

        # If the annotation is a link, extract the URL and its associated string
        if annotation_dict.get('/Subtype') == '/Link':
            url_dict = annotation_dict.get('/A')

            if url_dict is not None:
                url = url_dict.get('/URI')
                url_string = annotation_dict.get('/Contents')

                if url is not None:
                    # Check if the URL is working or broken
                    try:
                        response = requests.get(url)

                        if response.status_code == 200:
                            print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nWorking fine!")
                        else:
                            print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nBroken!")
                    except requests.exceptions.RequestException as e:
                        print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nBroken! Error: {e}")

# Close the PDF file
pdf_file.close()

运行后关联文本的输出结果为:

String - None

尝试过相关方案仍未解决问题,特此求助。


解决方案

使用PyMuPDF(fitz)库可以解决这个问题,它能获取链接的位置边界框,进而提取该区域内的关联文本。

1. 安装PyMuPDF

pip install pymupdf

2. 提取链接及关联文本的代码

import fitz
import requests

# 打开PDF文档
doc = fitz.open('abc.pdf')

for page_num in range(doc.page_count):
    page = doc[page_num]
    # 获取页面所有链接信息
    links = page.get_links()
    
    for link in links:
        # 仅处理外部URL链接
        if link.get('uri'):
            url = link['uri']
            # 将链接的坐标转为边界框对象
            rect = fitz.Rect(link['from'])
            # 提取边界框内的文本并去除多余空白
            link_text = page.get_textbox(rect).strip()
            
            # 检查URL可用性(可选步骤)
            status = "Working fine!"
            try:
                response = requests.get(url, timeout=5)
                if response.status_code != 200:
                    status = f"Broken! Status code: {response.status_code}"
            except requests.exceptions.RequestException as e:
                status = f"Broken! Error: {str(e)}"
            
            print(f"Page {page_num + 1}:")
            print(f"URL - {url}")
            print(f"关联文本 - {link_text if link_text else '未提取到文本'}")
            print(f"状态 - {status}\n")

doc.close()

代码说明

  • page.get_links():获取当前页面所有链接,包含链接的位置坐标和URI信息
  • fitz.Rect(link['from']):将链接的原始坐标转换为可操作的边界框对象
  • page.get_textbox(rect):精准提取边界框范围内的文本内容,strip()用于去除文本前后的空格和换行符

内容的提问来源于stack exchange,提问作者Sandy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 14:23:21