Ubuntu系统下Tesseract识别图片生成可见纯文本PDF的相关问题咨询
问题1:调整Tesseract配置使生成的PDF文字可见
问题原因
开启textonly_pdf参数后文字不可见,是因为未指定可见字体、或系统缺少对应字体依赖导致的。
解决步骤
- 先升级Tesseract到4.0以上版本,同时安装通用开源字体和对应语言识别包(以下示例包含简体中文支持,不需要可去掉
tesseract-ocr-chi-sim)
sudo apt update && sudo apt install --only-upgrade tesseract-ocr fonts-liberation tesseract-ocr-chi-sim
- 调整运行命令,指定PDF输出使用的可见字体即可:
# 纯英文识别场景 tesseract -c textonly_pdf=1 -c pdf_font_name="Liberation Sans" test.tif test pdf # 含简体中文识别场景,新增-l参数指定语言 tesseract -l chi_sim+eng -c textonly_pdf=1 -c pdf_font_name="Liberation Sans" test.tif test pdf
问题2:其他可实现需求的命令行/Python工具
命令行工具:ocrmypdf
是基于Tesseract封装的OCR专用工具,原生支持输出纯文本无原图的PDF,还自带倾斜校正、噪点清理等优化功能,识别准确率更高,Ubuntu下安装使用步骤:
- 安装
sudo apt install ocrmypdf tesseract-ocr-chi-sim
- 运行命令,
--output-type pdftext参数就是仅输出可编辑文本、不保留原始图片
ocrmypdf -l chi_sim+eng --output-type pdftext test.tif test_output.pdf
Python工具:pytesseract + reportlab
可以自定义文本排版、字体等细节,适合有定制化需求的场景:
- 安装依赖
sudo apt install python3-pip tesseract-ocr-chi-sim && pip3 install pytesseract reportlab
- 示例代码
import pytesseract from reportlab.pdfgen import canvas from reportlab.pdfbase import pdfmetrics from reportlab.pdfbase.ttfonts import TTFont # 注册中文字体(不需要中文支持可跳过这步,直接用默认Helvetica字体) pdfmetrics.registerFont(TTFont('LiberationSans', '/usr/share/fonts/truetype/liberation/LiberationSans-Regular.ttf')) # 识别图片文本 ocr_text = pytesseract.image_to_string('test.tif', lang='chi_sim+eng') # 生成纯文本PDF c = canvas.Canvas('output.pdf') c.setFont("LiberationSans", 12) current_y = 800 line_height = 16 for line in ocr_text.split('\n'): # 页面超出时自动换页 if current_y < 40: c.showPage() c.setFont("LiberationSans", 12) current_y = 800 if line.strip(): c.drawString(50, current_y, line.strip()) current_y -= line_height c.save()
内容的提问来源于stack exchange,提问作者Tinke
相关产品推荐
相关产品推荐

