You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

扫描PDF提取文本后特定内容过滤报错,求优化方案

问题分析

错误根源是filterAnanas接收了字符串列表作为匹配输入,但正则表达式的finditer方法要求传入单个字符串或字节对象,直接传列表就会触发类型错误。

优化方案

严格保持文本提取与内容过滤的职责分离:让get_text_from_image专注提取PDF各页文本并返回列表,filterAnanas专注处理单页文本的正则匹配,最后通过上层逻辑串联两个函数,遍历处理每一页内容。

修正后的代码示例

1. 文本提取函数(保留原有职责)

def get_text_from_image(pdf_path):
    # 原有OCR/PDF文本提取逻辑,返回各页文本组成的列表
    text_factuur_verdi = []
    # 此处省略具体的扫描PDF处理代码(如pytesseract、PyMuPDF调用)
    return text_factuur_verdi

2. 过滤函数(专注单字符串匹配)

import re

def filterAnanas(page_text):
    target_pattern = r"Ananas Crownless 14kg 10 Sweet CR Klasse I"
    # 查找所有匹配项并返回结果
    return [match.group() for match in re.finditer(target_pattern, page_text)]

3. 串联调用逻辑

# 提取所有页的文本
all_pages_text = get_text_from_image("你的扫描PDF路径.pdf")

# 遍历每一页执行过滤,收集所有匹配结果
all_matches = []
for page_content in all_pages_text:
    page_matches = filterAnanas(page_content)
    if page_matches:
        all_matches.extend(page_matches)

# 输出最终匹配结果
print(all_matches)

可选兼容优化(让过滤函数支持批量输入)

如果希望filterAnanas能直接接收列表输入,可在函数内部做兼容处理,核心匹配逻辑仍保持单字符串处理:

def filterAnanas(text_input):
    target_pattern = r"Ananas Crownless 14kg 10 Sweet CR Klasse I"
    all_matches = []
    
    if isinstance(text_input, list):
        # 处理列表输入,遍历每个元素
        for text in text_input:
            all_matches.extend([match.group() for match in re.finditer(target_pattern, text)])
    else:
        # 处理单个字符串输入
        all_matches = [match.group() for match in re.finditer(target_pattern, text_input)]
    
    return all_matches

# 直接调用即可
all_matches = filterAnanas(get_text_from_image("你的扫描PDF路径.pdf"))

内容的提问来源于stack exchange,提问作者mightycode Newton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 03:25:23