You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从PDF图片提取文本后筛选指定内容失败,求解决方法

问题:无法从PDF提取的文本中正确筛选目标字符串

提取到的文本内容

Koopliedenweg 38
Deb. nr. : 108636 2991 LN BARENDRECHT
Your VAT nr. : NL851703884B01 Nederland
Factuur datum : 19-11-21
Aantal Omschrijving Prijs Bedrag
Order number : 76372 Loading date : 15-11-21 Incoterm: : FOT
Your ref. : SCHOOLFRUIT Delivery date :
WK46
Verdi Import Schoolfruit
566 Ananas 
Crownless 14kg 10 Sweet CR Klasse I € 7,00 € 3.962,00
706 Appels Royal Gala 13kg 60/65 Generica PL Klasse I € 4,68 € 3.304,08
598 Peen Waspeen 14x1lkg 200-400 Generica BE Klasse I € 6,30 3.767,40
Order number : 76462 Loading date : 18-11-21 Incoterm: : FOT
Your ref. : SCHOOLFRUIT Delivery date

问题现象

尝试筛选字符串Appels Royal Gala 13kg时,最初代码返回空数组[];调整代码后返回异常结果['='],但测试其他普通文本时筛选功能正常。

代码尝试过程

最初代码及空数组输出

import io
from PIL import Image
import pytesseract
from wand.image import Image as wi

pdfFile = wi(filename = "C:\\Users\\engel\\Documents\\python\\docs\\fixedPDF.pdf", resolution = 300)
image = pdfFile.convert('jpeg')

imageBlobs = []

for img in image.sequence:
    imgPage = wi(image = img)
    imageBlobs.append(imgPage.make_blob('jpeg'))

extract = []

for imgBlob in imageBlobs:
    image = Image.open(io.BytesIO(imgBlob))
    text = pytesseract.image_to_string(image, lang = 'eng')
    extract.append(text)
    interested_string = 'Appels Royal Gala 13kg'
    line = [l[1] for l in extract if 'Appels Royal Gala 13kg' in l[1]]
print(line)

输出:[]

调整后代码及异常输出

仅修改筛选行:

line = [l[1] for l in extract if 'Appels Royal Gala 13kg' in l] 

输出:['=']

测试其他文本的正常结果

测试代码:

line = [l for l in extract if 'It was the best of' in l]

print(extract)

print(line)

输出:

['It was the best of\ntimes, it was the worst\nof times, it was the age\nof wisdom, it was the\nage of foolishness...\n\x0c']
['It was the best of\ntimes, it was the worst\nof times, it was the age\nof wisdom, it was the\nage of foolishness...\n\x0c']

解决方案及代码修改

问题根源

  1. extract中的元素是完整的页面文本字符串,而非列表。原代码中l[1]是取字符串的第二个字符,这就是返回['=']的原因——某页文本的第二个字符恰好是=。
  2. 筛选逻辑位置错误,原代码在循环内每次添加文本后立刻筛选,可能还未处理完所有页面。

修改后的代码

import io
from PIL import Image
import pytesseract
from wand.image import Image as wi

pdfFile = wi(filename = "C:\\Users\\engel\\Documents\\python\\docs\\fixedPDF.pdf", resolution = 300)
image = pdfFile.convert('jpeg')

imageBlobs = []

for img in image.sequence:
    imgPage = wi(image = img)
    imageBlobs.append(imgPage.make_blob('jpeg'))

extract = []
interested_string = 'Appels Royal Gala 13kg'
target_content = []

for imgBlob in imageBlobs:
    image = Image.open(io.BytesIO(imgBlob))
    # 加入荷兰语识别(文本含荷兰语词汇),提升OCR准确率
    text = pytesseract.image_to_string(image, lang='eng+nld')
    extract.append(text)
    # 按行拆分文本,逐行匹配目标字符串
    for line in text.split('\n'):
        if interested_string in line.strip():
            # 若只需目标字符串本身,替换为 target_content.append(interested_string)
            target_content.append(line.strip())

print(target_content)

额外优化建议

  • 若OCR识别存在字符/空格误差,可使用模糊匹配库(如fuzzywuzzy)降低匹配严格度;
  • 调整Tesseract的OCR参数(如添加荷兰语语言包),提升含非英语词汇文本的识别准确率。

内容的提问来源于stack exchange,提问作者mightycode Newton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 23:05:24