Python从PDF图片提取文本后筛选指定内容失败,求解决方法
问题:无法从PDF提取的文本中正确筛选目标字符串
提取到的文本内容
Koopliedenweg 38 Deb. nr. : 108636 2991 LN BARENDRECHT Your VAT nr. : NL851703884B01 Nederland Factuur datum : 19-11-21 Aantal Omschrijving Prijs Bedrag Order number : 76372 Loading date : 15-11-21 Incoterm: : FOT Your ref. : SCHOOLFRUIT Delivery date : WK46 Verdi Import Schoolfruit 566 Ananas Crownless 14kg 10 Sweet CR Klasse I € 7,00 € 3.962,00 706 Appels Royal Gala 13kg 60/65 Generica PL Klasse I € 4,68 € 3.304,08 598 Peen Waspeen 14x1lkg 200-400 Generica BE Klasse I € 6,30 3.767,40 Order number : 76462 Loading date : 18-11-21 Incoterm: : FOT Your ref. : SCHOOLFRUIT Delivery date
问题现象
尝试筛选字符串Appels Royal Gala 13kg时,最初代码返回空数组[];调整代码后返回异常结果['='],但测试其他普通文本时筛选功能正常。
代码尝试过程
最初代码及空数组输出
import io from PIL import Image import pytesseract from wand.image import Image as wi pdfFile = wi(filename = "C:\\Users\\engel\\Documents\\python\\docs\\fixedPDF.pdf", resolution = 300) image = pdfFile.convert('jpeg') imageBlobs = [] for img in image.sequence: imgPage = wi(image = img) imageBlobs.append(imgPage.make_blob('jpeg')) extract = [] for imgBlob in imageBlobs: image = Image.open(io.BytesIO(imgBlob)) text = pytesseract.image_to_string(image, lang = 'eng') extract.append(text) interested_string = 'Appels Royal Gala 13kg' line = [l[1] for l in extract if 'Appels Royal Gala 13kg' in l[1]] print(line)
输出:[]
调整后代码及异常输出
仅修改筛选行:
line = [l[1] for l in extract if 'Appels Royal Gala 13kg' in l]
输出:['=']
测试其他文本的正常结果
测试代码:
line = [l for l in extract if 'It was the best of' in l] print(extract) print(line)
输出:
['It was the best of\ntimes, it was the worst\nof times, it was the age\nof wisdom, it was the\nage of foolishness...\n\x0c'] ['It was the best of\ntimes, it was the worst\nof times, it was the age\nof wisdom, it was the\nage of foolishness...\n\x0c']
解决方案及代码修改
问题根源
extract中的元素是完整的页面文本字符串,而非列表。原代码中l[1]是取字符串的第二个字符,这就是返回['=']的原因——某页文本的第二个字符恰好是=。- 筛选逻辑位置错误,原代码在循环内每次添加文本后立刻筛选,可能还未处理完所有页面。
修改后的代码
import io from PIL import Image import pytesseract from wand.image import Image as wi pdfFile = wi(filename = "C:\\Users\\engel\\Documents\\python\\docs\\fixedPDF.pdf", resolution = 300) image = pdfFile.convert('jpeg') imageBlobs = [] for img in image.sequence: imgPage = wi(image = img) imageBlobs.append(imgPage.make_blob('jpeg')) extract = [] interested_string = 'Appels Royal Gala 13kg' target_content = [] for imgBlob in imageBlobs: image = Image.open(io.BytesIO(imgBlob)) # 加入荷兰语识别(文本含荷兰语词汇),提升OCR准确率 text = pytesseract.image_to_string(image, lang='eng+nld') extract.append(text) # 按行拆分文本,逐行匹配目标字符串 for line in text.split('\n'): if interested_string in line.strip(): # 若只需目标字符串本身,替换为 target_content.append(interested_string) target_content.append(line.strip()) print(target_content)
额外优化建议
- 若OCR识别存在字符/空格误差,可使用模糊匹配库(如
fuzzywuzzy)降低匹配严格度; - 调整Tesseract的OCR参数(如添加荷兰语语言包),提升含非英语词汇文本的识别准确率。
内容的提问来源于stack exchange,提问作者mightycode Newton
相关产品推荐
相关产品推荐

