Django5 + Python3.12实现PDF上传处理转Excel并自动下载求助
实现PDF转Excel并自动下载的Django5应用代码
1. 安装依赖
首先安装所需的Python包:
pip install django==5.0 PyPDF2 openpyxl
2. Django项目配置
假设你的Django项目名为pdf_converter,应用名为pdf2excel:
2.1 配置settings.py
在settings.py中添加以下配置:
INSTALLED_APPS = [ # 其他默认应用 'pdf2excel', ] # 临时文件存储配置(用于临时保存上传的PDF) MEDIA_ROOT = BASE_DIR / 'media' MEDIA_URL = '/media/'
3. 视图函数(views.py)
在pdf2excel/views.py中编写核心逻辑:
import os from django.shortcuts import render from django.http import HttpResponse from PyPDF2 import PdfReader from openpyxl import Workbook from django.conf import settings from django.core.files.storage import default_storage def pdf_to_excel(request): if request.method == 'POST': # 获取上传的PDF文件 pdf_file = request.FILES.get('pdf_file') if not pdf_file: return render(request, 'upload.html', {'error': '请选择PDF文件'}) # 临时保存PDF文件 temp_pdf_path = os.path.join(settings.MEDIA_ROOT, pdf_file.name) with default_storage.open(temp_pdf_path, 'wb+') as destination: for chunk in pdf_file.chunks(): destination.write(chunk) try: # 读取PDF内容 reader = PdfReader(temp_pdf_path) pdf_content = [] for page in reader.pages: text = page.extract_text() if text: # 按行拆分文本,可根据业务需求自定义处理逻辑 lines = text.split('\n') pdf_content.extend([line.strip() for line in lines if line.strip()]) # 创建Excel文件并写入数据 wb = Workbook() ws = wb.active ws.title = 'PDF内容' # 写入表头 ws.append(['序号', '内容']) for idx, content in enumerate(pdf_content, 1): ws.append([idx, content]) # 生成下载响应 response = HttpResponse(content_type='application/vnd.openxmlformats-officedocument.spreadsheetml.sheet') response['Content-Disposition'] = f'attachment; filename="pdf_converted.xlsx"' wb.save(response) return response except Exception as e: return render(request, 'upload.html', {'error': f'处理失败:{str(e)}'}) finally: # 清理临时PDF文件 if os.path.exists(temp_pdf_path): os.remove(temp_pdf_path) # GET请求返回上传页面 return render(request, 'upload.html')
4. 上传模板(templates/upload.html)
在pdf2excel/templates/upload.html中创建上传表单:
<!DOCTYPE html> <html> <head> <title>PDF转Excel</title> </head> <body> <h1>上传PDF文件转Excel</h1> {% if error %} <p style="color: red;">{{ error }}</p> {% endif %} <form method="post" enctype="multipart/form-data"> {% csrf_token %} <input type="file" name="pdf_file" accept=".pdf" required> <button type="submit">转换并下载</button> </form> </body> </html>
5. URL配置(urls.py)
项目根urls.py
from django.contrib import admin from django.urls import path, include from django.conf import settings from django.conf.urls.static import static urlpatterns = [ path('admin/', admin.site.urls), path('', include('pdf2excel.urls')), ] + static(settings.MEDIA_URL, document_root=settings.MEDIA_ROOT)
应用urls.py(pdf2excel/urls.py)
from django.urls import path from . import views urlpatterns = [ path('', views.pdf_to_excel, name='pdf_to_excel'), ]
注意事项
- PDF内容处理:示例仅做基础文本提取,若需处理表格或结构化内容,可替换
PyPDF2为pdfplumber以获得更精准的格式解析。 - 大文件优化:处理大体积PDF时,建议采用流式读取+临时文件分片处理,避免内存溢出。
- 异常扩展:可针对性捕获PDF格式错误、文件损坏等特定异常,优化用户提示。
内容的提问来源于stack exchange,提问作者Satish Kumar
相关产品推荐
相关产品推荐

