Flask中PDF上传与主题(含标题及描述)提取实现求助
Flask PDF主题提取实现方案
1. 搭建基础上传路由
先写核心的Flask上传接口,只接受PDF文件,处理上传后调用提取逻辑:
from flask import Flask, request, jsonify import os from werkzeug.utils import secure_filename app = Flask(__name__) app.config['UPLOAD_FOLDER'] = './uploads' app.config['ALLOWED_EXTENSIONS'] = {'pdf'} def allowed_file(filename): return '.' in filename and filename.rsplit('.', 1)[1].lower() in app.config['ALLOWED_EXTENSIONS'] @app.route('/upload-pdf', methods=['POST']) def upload_pdf(): if 'file' not in request.files: return jsonify({'error': '未上传文件'}), 400 file = request.files['file'] if file.filename == '': return jsonify({'error': '未选择文件'}), 400 if file and allowed_file(file.filename): filename = secure_filename(file.filename) file_path = os.path.join(app.config['UPLOAD_FOLDER'], filename) file.save(file_path) # 调用主题提取函数 topics = extract_topics_from_pdf(file_path) # 可选:删除临时文件 os.remove(file_path) return jsonify({'topics': topics}) return jsonify({'error': '仅支持PDF格式文件'}), 400 if __name__ == '__main__': os.makedirs(app.config['UPLOAD_FOLDER'], exist_ok=True) app.run(debug=True)
2. 核心:PDF标题与描述关联提取
用pdfplumber工具(能获取字体、字号信息,比单纯提取文本更精准),通过字体特征区分标题和正文:
先安装依赖:pip install pdfplumber flask
提取函数代码:
import pdfplumber def extract_topics_from_pdf(file_path): topics = [] current_title = None current_desc = [] with pdfplumber.open(file_path) as pdf: for page in pdf.pages: # 获取带字体属性的文本片段 words = page.extract_words(extra_attrs=['fontname', 'size']) for word in words: # 自定义标题判断规则:字号>14且字体含加粗标识(根据你的PDF调整参数) is_title = word['size'] > 14 and ('Bold' in word['fontname'] or 'bold' in word['fontname']) if is_title: # 先保存上一个主题的内容 if current_title and current_desc: topics.append({ 'title': current_title.strip(), 'desc': ' '.join(current_desc).strip() }) current_title = word['text'] current_desc = [] else: # 只有存在当前标题时,才收集描述内容 if current_title: current_desc.append(word['text']) # 处理最后一个主题 if current_title and current_desc: topics.append({ 'title': current_title.strip(), 'desc': ' '.join(current_desc).strip() }) return topics
3. 适配特殊PDF的补充方案
扫描版PDF(图片转文本)
如果是扫描生成的PDF,需要先做OCR识别,再提取主题:
安装依赖:pip install pdf2image pytesseract
from pdf2image import convert_from_path import pytesseract import re def extract_scanned_pdf(file_path): images = convert_from_path(file_path) full_text = '' for img in images: full_text += pytesseract.image_to_string(img) # 假设标题是"数字. 主题名"格式,用正则匹配分割 topics = [] pattern = re.compile(r'(\d+\. .+?)(?=\n\d+\. |\Z)', re.DOTALL) for match in pattern.finditer(full_text): content = match.group(1).split('\n', 1) if len(content) >=2: topics.append({'title': content[0].strip(), 'desc': content[1].strip()}) return topics
结构化PDF(带大纲)
如果PDF有内置大纲,可以直接用大纲标题匹配对应页面内容:
def extract_outline_pdf(file_path): topics = [] with pdfplumber.open(file_path) as pdf: outline = pdf.outline for item in outline: # 获取大纲对应的页码,提取该页及后续内容直到下一个大纲 page_num = item.page_number - 1 # pdfplumber页码从0开始 page = pdf.pages[page_num] text = page.extract_text() # 这里可以根据实际情况切割标题和描述 if text: parts = text.split('\n', 1) if len(parts)>=2: topics.append({'title': parts[0].strip(), 'desc': parts[1].strip()}) return topics
4. 实用调整建议
- 测试你的目标PDF,修改标题判断的字号、字体规则,不同PDF格式差异大,需要针对性调整。
- 大文件可以分页处理,避免内存占用过高。
- 加入异常捕获,比如处理损坏的PDF文件,返回明确错误信息。
内容的提问来源于stack exchange,提问作者Ahsan Hafeez
相关产品推荐
相关产品推荐

