基于Tesseract-OCR的孟加拉语PDF转DOCX转换器优化求助
孟加拉语PDF转换器优化方案:文本、表格及格式精准提取
现有代码核心问题
- 错误将DOCX文件传入
pytesseract.image_to_string():Tesseract仅支持图像格式(PNG/JPG等),无法直接识别DOCX pdf2docx对非拉丁文字符(如孟加拉语)原生支持有限,转档易丢失字符或格式- 未区分原生可编辑PDF和扫描版PDF,两者处理逻辑完全不同
一、核心优化思路
- 区分PDF类型
- 原生PDF:直接提取结构化文本、表格和格式,无需OCR
- 扫描版PDF:先转成图像,再用Tesseract OCR识别,同时保留布局信息
- 工具选型调整
- 原生PDF:用
PyMuPDF提取带格式文本,pdfplumber提取表格 - 扫描版PDF:用
pdf2image转图像,配合Tesseract孟加拉语训练库 - DOCX输出:用
python-docx手动构建文档,保留对齐、制表符等格式
- 原生PDF:用
二、优化后的代码实现
1. 依赖安装
pip install pymupdf pdfplumber pdf2image python-docx pytesseract
注意:需确保Tesseract已安装孟加拉语语言包,可通过tesseract --list-langs验证是否包含ben
2. 原生PDF处理(带格式提取)
import os from pathlib import Path import fitz # PyMuPDF import pdfplumber from docx import Document from docx.shared import Pt from docx.enum.text import WD_ALIGN_PARAGRAPH from tkinter import messagebox as mb current_path = os.path.dirname(os.path.abspath(__file__)) input_pdf = os.path.join(current_path, "input.pdf") output_docx = os.path.join(current_path, "output.docx") def process_native_pdf(pdf_path, docx_path): doc = Document() # 提取文本及格式 with fitz.open(pdf_path) as pdf: for page in pdf: blocks = page.get_text("dict")["blocks"] for block in blocks: if block["type"] == 0: # 文本块 para = doc.add_paragraph() for line in block["lines"]: for span in line["spans"]: run = para.add_run(span["text"]) # 设置孟加拉语字体(需系统已安装,如Siyam Rupali) run.font.name = "Siyam Rupali" run.font.size = Pt(span["size"]) # 匹配原对齐方式 if span["flags"] & 2: para.alignment = WD_ALIGN_PARAGRAPH.CENTER elif span["flags"] & 4: para.alignment = WD_ALIGN_PARAGRAPH.RIGHT else: para.alignment = WD_ALIGN_PARAGRAPH.LEFT # 提取表格并插入DOCX with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: tables = page.extract_tables() for table in tables: doc_table = doc.add_table(rows=len(table), cols=len(table[0])) for i, row in enumerate(table): for j, cell in enumerate(row): doc_table.cell(i,j).text = cell if cell else "" doc.save(docx_path) print("原生PDF处理完成,已保存为DOCX") # 执行处理 if Path(input_pdf).is_file() and input_pdf.lower().endswith(".pdf"): try: process_native_pdf(input_pdf, output_docx) except Exception as ex: mb.showerror("错误", f"处理失败:{str(ex)}") else: mb.showerror("提示", "未找到PDF文件!")
3. 扫描版PDF处理(OCR+格式保留)
import os from pathlib import Path import pdf2image import pytesseract from docx import Document from tkinter import messagebox as mb current_path = os.path.dirname(os.path.abspath(__file__)) input_pdf = os.path.join(current_path, "input.pdf") output_docx = os.path.join(current_path, "output_ocr.docx") # 若Tesseract未加入系统环境变量,需指定路径 # pytesseract.pytesseract.tesseract_cmd = r"C:\Program Files\Tesseract-OCR\tesseract.exe" def process_scanned_pdf(pdf_path, docx_path): doc = Document() # 将PDF转成图像 images = pdf2image.convert_from_path(pdf_path) for img in images: # 用Tesseract识别,保留段落布局 text = pytesseract.image_to_string(img, lang="ben", config="--psm 6") para = doc.add_paragraph(text) # 设置孟加拉语字体 for run in para.runs: run.font.name = "Siyam Rupali" doc.save(docx_path) print("扫描版PDF OCR处理完成,已保存为DOCX") # 执行处理 if Path(input_pdf).is_file() and input_pdf.lower().endswith(".pdf"): try: process_scanned_pdf(input_pdf, output_docx) except Exception as ex: mb.showerror("错误", f"处理失败:{str(ex)}") else: mb.showerror("提示", "未找到PDF文件!")
三、关键注意事项
- 确保系统安装孟加拉语字体(如Siyam Rupali、SolaimanLipi),避免DOCX乱码
- Tesseract需安装最新版,且将
ben.traineddata放在Tesseract的tessdata目录下 - 混合类型PDF(部分原生、部分扫描)可结合两种逻辑,先尝试原生提取,失败自动切换OCR
内容的提问来源于stack exchange,提问作者Mohammad Yasin
相关产品推荐
相关产品推荐

