如何将PDF转换为可搜索的结构化JSON表格?技术求助
技术解决方案
一、PDF内容导出(从Archive.org获取文本)
1. Python工具库直接抓取+提取文本
Archive.org的PDF可通过页面下载按钮获取原始文件链接,用pdfplumber(文本提取精度优于PyPDF2)配合requests实现下载与文本提取:
- 安装依赖:
pip install pdfplumber requests
- 代码示例:
import requests import pdfplumber # 替换为目标PDF的原始下载链接(从Archive.org页面的「下载」选项获取) pdf_url = "https://archive.org/download/isbn_9781585746613/isbn_9781585746613.pdf" response = requests.get(pdf_url, headers={"User-Agent": "Mozilla/5.0"}) with open("target.pdf", "wb") as f: f.write(response.content) # 提取每页文本并按页码存储 page_texts = [] with pdfplumber.open("target.pdf") as pdf: for page in pdf.pages: extracted_text = page.extract_text() if extracted_text: page_texts.append({ "page_number": page.page_number, "raw_content": extracted_text.strip() })
2. Archive.org API辅助获取
若直接下载受限,可先通过API获取PDF元数据中的原始文件链接:
import requests identifier = "isbn_9781585746613" api_url = f"https://archive.org/metadata/{identifier}" response = requests.get(api_url).json() # 从元数据中筛选PDF格式的下载链接 pdf_download_url = None for file in response["files"]: if file["format"] == "PDF" and file["name"].endswith(".pdf"): pdf_download_url = f"https://archive.org/download/{identifier}/{file['name']}" break # 后续下载与文本提取逻辑同方法1
二、页面板块拆分
板块拆分需根据PDF排版类型选择方案:
1. 结构化PDF(Word生成、有明确层级)
利用pdfplumber的extract_words()获取文本块的字体大小、位置信息,区分标题与内容板块:
with pdfplumber.open("target.pdf") as pdf: target_page = pdf.pages[7] # 示例第8页对应索引7 words = target_page.extract_words() sections = [] current_section = {"title": "", "content": ""} # 假设标题字体大于12号,可根据实际PDF调整阈值 title_font_threshold = 12 for word in words: if word["size"] > title_font_threshold: # 保存上一个板块 if current_section["title"]: sections.append(current_section) current_section = {"title": word["text"], "content": word["text"]} else: current_section["content"] += " " + word["text"] # 保存最后一个板块 sections.append(current_section) # sections即为该页面拆分后的板块列表
2. 非结构化PDF(扫描件、无规则排版)
需先通过OCR识别图片文本,推荐用pdf2image+pytesseract:
- 安装依赖:
pip install pdf2image pytesseract # 额外安装Tesseract引擎:Ubuntu用sudo apt install tesseract-ocr,Windows需下载安装包并配置环境变量
- 代码示例:
from pdf2image import convert_from_path import pytesseract # 转换PDF为图片 pages = convert_from_path("target.pdf", dpi=300) # 提高dpi提升识别精度 for page_idx, page_img in enumerate(pages): # OCR识别文本 ocr_text = pytesseract.image_to_string(page_img, lang="eng") # 可通过关键词(如Chapter、Section)拆分板块 sections = [s.strip() for s in ocr_text.split("\n\n") if s.strip()] # 存储板块数据 page_data = { "page_number": page_idx + 1, "sections": sections }
三、整合至JSON页面
1. 生成静态JSON文件
将所有页面的板块数据整理为标准JSON格式:
import json # 假设all_page_sections是所有页面的板块数据列表 output_data = { "pdf_identifier": "isbn_9781585746613", "total_pages": len(all_page_sections), "content": all_page_sections } with open("pdf_content.json", "w", encoding="utf-8") as f: json.dump(output_data, f, indent=2, ensure_ascii=False)
2. 构建JSON可视化页面
用简单HTML加载并展示JSON数据:
<!DOCTYPE html> <html> <head> <title>PDF Content Viewer</title> <style> #content-container { margin: 20px; font-family: Arial, sans-serif; } .page-section { margin: 15px 0; padding: 10px; border: 1px solid #eee; border-radius: 4px; } .page-number { font-weight: bold; color: #2c3e50; } .section-title { font-size: 1.2em; margin: 5px 0; } .section-content { color: #34495e; line-height: 1.5; } </style> </head> <body> <div id="content-container"></div> <script> fetch('pdf_content.json') .then(res => res.json()) .then(data => { const container = document.getElementById('content-container'); data.content.forEach(page => { const pageDiv = document.createElement('div'); pageDiv.innerHTML = `<div class="page-number">Page ${page.page_number}</div>`; page.sections.forEach(section => { const sectionDiv = document.createElement('div'); sectionDiv.className = 'page-section'; sectionDiv.innerHTML = ` <div class="section-title">${section.title}</div> <div class="section-content">${section.content.replace(/\n/g, '<br>')}</div> `; pageDiv.appendChild(sectionDiv); }); container.appendChild(pageDiv); }); }); </script> </body> </html>
常见问题处理
- Archive.org请求限制:添加
User-Agent请求头模拟浏览器,或设置请求间隔避免IP被封。 - 文本提取混乱:复杂排版PDF可单独用
pdfplumber.extract_tables()提取表格,单独处理。 - OCR识别精度低:添加对应语言包(如中文
chi_sim),或对图片做对比度、分辨率预处理。
内容的提问来源于stack exchange,提问作者TexanBuilder1 Studio
相关产品推荐
相关产品推荐

