You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将PDF转换为可搜索的结构化JSON表格?技术求助

技术解决方案

一、PDF内容导出(从Archive.org获取文本)

1. Python工具库直接抓取+提取文本

Archive.org的PDF可通过页面下载按钮获取原始文件链接,用pdfplumber(文本提取精度优于PyPDF2)配合requests实现下载与文本提取:

  • 安装依赖:
pip install pdfplumber requests
  • 代码示例:
import requests
import pdfplumber

# 替换为目标PDF的原始下载链接(从Archive.org页面的「下载」选项获取)
pdf_url = "https://archive.org/download/isbn_9781585746613/isbn_9781585746613.pdf"
response = requests.get(pdf_url, headers={"User-Agent": "Mozilla/5.0"})

with open("target.pdf", "wb") as f:
    f.write(response.content)

# 提取每页文本并按页码存储
page_texts = []
with pdfplumber.open("target.pdf") as pdf:
    for page in pdf.pages:
        extracted_text = page.extract_text()
        if extracted_text:
            page_texts.append({
                "page_number": page.page_number,
                "raw_content": extracted_text.strip()
            })

2. Archive.org API辅助获取

若直接下载受限,可先通过API获取PDF元数据中的原始文件链接:

import requests

identifier = "isbn_9781585746613"
api_url = f"https://archive.org/metadata/{identifier}"
response = requests.get(api_url).json()

# 从元数据中筛选PDF格式的下载链接
pdf_download_url = None
for file in response["files"]:
    if file["format"] == "PDF" and file["name"].endswith(".pdf"):
        pdf_download_url = f"https://archive.org/download/{identifier}/{file['name']}"
        break

# 后续下载与文本提取逻辑同方法1

二、页面板块拆分

板块拆分需根据PDF排版类型选择方案:

1. 结构化PDF(Word生成、有明确层级)

利用pdfplumber的extract_words()获取文本块的字体大小、位置信息,区分标题与内容板块:

with pdfplumber.open("target.pdf") as pdf:
    target_page = pdf.pages[7]  # 示例第8页对应索引7
    words = target_page.extract_words()
    
    sections = []
    current_section = {"title": "", "content": ""}
    # 假设标题字体大于12号,可根据实际PDF调整阈值
    title_font_threshold = 12

    for word in words:
        if word["size"] > title_font_threshold:
            # 保存上一个板块
            if current_section["title"]:
                sections.append(current_section)
            current_section = {"title": word["text"], "content": word["text"]}
        else:
            current_section["content"] += " " + word["text"]
    # 保存最后一个板块
    sections.append(current_section)

# sections即为该页面拆分后的板块列表

2. 非结构化PDF(扫描件、无规则排版)

需先通过OCR识别图片文本,推荐用pdf2image+pytesseract:

  • 安装依赖:
pip install pdf2image pytesseract
# 额外安装Tesseract引擎:Ubuntu用sudo apt install tesseract-ocr,Windows需下载安装包并配置环境变量
  • 代码示例:
from pdf2image import convert_from_path
import pytesseract

# 转换PDF为图片
pages = convert_from_path("target.pdf", dpi=300)  # 提高dpi提升识别精度

for page_idx, page_img in enumerate(pages):
    # OCR识别文本
    ocr_text = pytesseract.image_to_string(page_img, lang="eng")
    # 可通过关键词(如Chapter、Section)拆分板块
    sections = [s.strip() for s in ocr_text.split("\n\n") if s.strip()]
    # 存储板块数据
    page_data = {
        "page_number": page_idx + 1,
        "sections": sections
    }

三、整合至JSON页面

1. 生成静态JSON文件

将所有页面的板块数据整理为标准JSON格式:

import json

# 假设all_page_sections是所有页面的板块数据列表
output_data = {
    "pdf_identifier": "isbn_9781585746613",
    "total_pages": len(all_page_sections),
    "content": all_page_sections
}

with open("pdf_content.json", "w", encoding="utf-8") as f:
    json.dump(output_data, f, indent=2, ensure_ascii=False)

2. 构建JSON可视化页面

用简单HTML加载并展示JSON数据:

<!DOCTYPE html>
<html>
<head>
    <title>PDF Content Viewer</title>
    <style>
        #content-container { margin: 20px; font-family: Arial, sans-serif; }
        .page-section { margin: 15px 0; padding: 10px; border: 1px solid #eee; border-radius: 4px; }
        .page-number { font-weight: bold; color: #2c3e50; }
        .section-title { font-size: 1.2em; margin: 5px 0; }
        .section-content { color: #34495e; line-height: 1.5; }
    </style>
</head>
<body>
    <div id="content-container"></div>
    <script>
        fetch('pdf_content.json')
            .then(res => res.json())
            .then(data => {
                const container = document.getElementById('content-container');
                data.content.forEach(page => {
                    const pageDiv = document.createElement('div');
                    pageDiv.innerHTML = `<div class="page-number">Page ${page.page_number}</div>`;
                    page.sections.forEach(section => {
                        const sectionDiv = document.createElement('div');
                        sectionDiv.className = 'page-section';
                        sectionDiv.innerHTML = `
                            <div class="section-title">${section.title}</div>
                            <div class="section-content">${section.content.replace(/\n/g, '<br>')}</div>
                        `;
                        pageDiv.appendChild(sectionDiv);
                    });
                    container.appendChild(pageDiv);
                });
            });
    </script>
</body>
</html>

常见问题处理

  • Archive.org请求限制:添加User-Agent请求头模拟浏览器,或设置请求间隔避免IP被封。
  • 文本提取混乱:复杂排版PDF可单独用pdfplumber.extract_tables()提取表格,单独处理。
  • OCR识别精度低:添加对应语言包(如中文chi_sim),或对图片做对比度、分辨率预处理。

内容的提问来源于stack exchange,提问作者TexanBuilder1 Studio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 00:37:35