如何用PyPDF2/PyPDF4提取PDF中URL关联的文本内容?
问题描述
我有一份包含超链接的PDF文档,需要提取URL对应的关联文本。使用PyPDF2和PyPDF4库时,仅能成功提取URL,无法获取绑定的文本内容——比如PDF里有一段绑定了URL的文本“Check this link out”,只能提取到URL,却拿不到这段关联文本。
当前使用的PyPDF4代码如下:
import PyPDF4 import requests # Open the PDF file pdf_file = open('abc.pdf', 'rb') # Create a PDF reader object pdf_reader = PyPDF4.PdfFileReader(pdf_file) # Loop through each page of the PDF for page_num in range(len(pdf_reader.pages)): # Get the page object page = pdf_reader.pages[page_num] # Extract the annotations from the page annotations = page.get('/Annots') # If there are no annotations, skip to the next page if not annotations: continue # Loop through each annotation for annotation in annotations: # Get the annotation dictionary annotation_dict = annotation.getObject() # If the annotation is a link, extract the URL and its associated string if annotation_dict.get('/Subtype') == '/Link': url_dict = annotation_dict.get('/A') if url_dict is not None: url = url_dict.get('/URI') url_string = annotation_dict.get('/Contents') if url is not None: # Check if the URL is working or broken try: response = requests.get(url) if response.status_code == 200: print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nWorking fine!") else: print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nBroken!") except requests.exceptions.RequestException as e: print(f"Page {page_num + 1}: URL - {url}\nString - {url_string}\nBroken! Error: {e}") # Close the PDF file pdf_file.close()
运行后关联文本的输出结果为:
String - None
尝试过相关方案仍未解决问题,特此求助。
解决方案
使用PyMuPDF(fitz)库可以解决这个问题,它能获取链接的位置边界框,进而提取该区域内的关联文本。
1. 安装PyMuPDF
pip install pymupdf
2. 提取链接及关联文本的代码
import fitz import requests # 打开PDF文档 doc = fitz.open('abc.pdf') for page_num in range(doc.page_count): page = doc[page_num] # 获取页面所有链接信息 links = page.get_links() for link in links: # 仅处理外部URL链接 if link.get('uri'): url = link['uri'] # 将链接的坐标转为边界框对象 rect = fitz.Rect(link['from']) # 提取边界框内的文本并去除多余空白 link_text = page.get_textbox(rect).strip() # 检查URL可用性(可选步骤) status = "Working fine!" try: response = requests.get(url, timeout=5) if response.status_code != 200: status = f"Broken! Status code: {response.status_code}" except requests.exceptions.RequestException as e: status = f"Broken! Error: {str(e)}" print(f"Page {page_num + 1}:") print(f"URL - {url}") print(f"关联文本 - {link_text if link_text else '未提取到文本'}") print(f"状态 - {status}\n") doc.close()
代码说明
page.get_links():获取当前页面所有链接,包含链接的位置坐标和URI信息fitz.Rect(link['from']):将链接的原始坐标转换为可操作的边界框对象page.get_textbox(rect):精准提取边界框范围内的文本内容,strip()用于去除文本前后的空格和换行符
内容的提问来源于stack exchange,提问作者Sandy
相关产品推荐
相关产品推荐

