You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Django5 + Python3.12实现PDF上传处理转Excel并自动下载求助

实现PDF转Excel并自动下载的Django5应用代码

1. 安装依赖

首先安装所需的Python包:

pip install django==5.0 PyPDF2 openpyxl

2. Django项目配置

假设你的Django项目名为pdf_converter,应用名为pdf2excel:

2.1 配置settings.py

在settings.py中添加以下配置:

INSTALLED_APPS = [
    # 其他默认应用
    'pdf2excel',
]

# 临时文件存储配置(用于临时保存上传的PDF)
MEDIA_ROOT = BASE_DIR / 'media'
MEDIA_URL = '/media/'

3. 视图函数(views.py)

在pdf2excel/views.py中编写核心逻辑:

import os
from django.shortcuts import render
from django.http import HttpResponse
from PyPDF2 import PdfReader
from openpyxl import Workbook
from django.conf import settings
from django.core.files.storage import default_storage

def pdf_to_excel(request):
    if request.method == 'POST':
        # 获取上传的PDF文件
        pdf_file = request.FILES.get('pdf_file')
        if not pdf_file:
            return render(request, 'upload.html', {'error': '请选择PDF文件'})
        
        # 临时保存PDF文件
        temp_pdf_path = os.path.join(settings.MEDIA_ROOT, pdf_file.name)
        with default_storage.open(temp_pdf_path, 'wb+') as destination:
            for chunk in pdf_file.chunks():
                destination.write(chunk)
        
        try:
            # 读取PDF内容
            reader = PdfReader(temp_pdf_path)
            pdf_content = []
            for page in reader.pages:
                text = page.extract_text()
                if text:
                    # 按行拆分文本,可根据业务需求自定义处理逻辑
                    lines = text.split('\n')
                    pdf_content.extend([line.strip() for line in lines if line.strip()])
            
            # 创建Excel文件并写入数据
            wb = Workbook()
            ws = wb.active
            ws.title = 'PDF内容'
            # 写入表头
            ws.append(['序号', '内容'])
            for idx, content in enumerate(pdf_content, 1):
                ws.append([idx, content])
            
            # 生成下载响应
            response = HttpResponse(content_type='application/vnd.openxmlformats-officedocument.spreadsheetml.sheet')
            response['Content-Disposition'] = f'attachment; filename="pdf_converted.xlsx"'
            wb.save(response)
            
            return response
        except Exception as e:
            return render(request, 'upload.html', {'error': f'处理失败:{str(e)}'})
        finally:
            # 清理临时PDF文件
            if os.path.exists(temp_pdf_path):
                os.remove(temp_pdf_path)
    
    # GET请求返回上传页面
    return render(request, 'upload.html')

4. 上传模板(templates/upload.html)

在pdf2excel/templates/upload.html中创建上传表单:

<!DOCTYPE html>
<html>
<head>
    <title>PDF转Excel</title>
</head>
<body>
    <h1>上传PDF文件转Excel</h1>
    {% if error %}
        <p style="color: red;">{{ error }}</p>
    {% endif %}
    <form method="post" enctype="multipart/form-data">
        {% csrf_token %}
        <input type="file" name="pdf_file" accept=".pdf" required>
        <button type="submit">转换并下载</button>
    </form>
</body>
</html>

5. URL配置(urls.py)

项目根urls.py

from django.contrib import admin
from django.urls import path, include
from django.conf import settings
from django.conf.urls.static import static

urlpatterns = [
    path('admin/', admin.site.urls),
    path('', include('pdf2excel.urls')),
] + static(settings.MEDIA_URL, document_root=settings.MEDIA_ROOT)

应用urls.py(pdf2excel/urls.py)

from django.urls import path
from . import views

urlpatterns = [
    path('', views.pdf_to_excel, name='pdf_to_excel'),
]

注意事项

  • PDF内容处理:示例仅做基础文本提取,若需处理表格或结构化内容,可替换PyPDF2为pdfplumber以获得更精准的格式解析。
  • 大文件优化:处理大体积PDF时,建议采用流式读取+临时文件分片处理,避免内存溢出。
  • 异常扩展:可针对性捕获PDF格式错误、文件损坏等特定异常,优化用户提示。

内容的提问来源于stack exchange,提问作者Satish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 14:42:50