You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Lambda函数中批量传递S3文件至AWS Textract API提速咨询

问题与优化方案

问题背景

我有一张包含多文本字段的PDF图片,已裁剪出对应每个文本框的小图,用字典管理各图片对应的字段。原代码通过串行调用AWS Textract API提取文本,处理数百张图耗时约7分钟,需要缩短至45秒内。尝试过传入S3文件名列表但报错,不清楚如何实现并行处理或批量提交并保持字段对应关系。

原字典结构示例:

aws_output = {
    "Section A": {"1": extract_text(imgs['Section A']['1'], "sa1"), 
                  "2": extract_text(imgs['Section A']['2'], "sa2")}, 
    "Section B": {"1": extract_text(imgs['Section B']['1'], "sb1")}
}

原提取函数:

def extract_text(img, loc):
        img_obj = Image.fromarray(img).convert('RGB')
        out_img_obj = io.BytesIO()
        img_obj.save(out_img_obj, format="png")
        out_img_obj.seek(0)
        file_name = key_id + "_" + loc + ".png"
        s3.Bucket(bucket_name).put_object(Key=file_name, Body=out_img_obj, ContentType="image/png")
        response = textract_client.detect_document_text(
            Document={
                'S3Object': {
                    'Bucket': bucket_name,
                    'Name': file_name
                }
            }
        )
        status_code = response['ResponseMetadata']['HTTPStatusCode']
        try: 
            status_code == 200
            text_len = {}
            for y in range(len(response['Blocks'])):
                if 'Text' in response['Blocks'][y]:
                    text_len[y] = len(response['Blocks'][y]['Text'])
                else:
                    pass
            if bool(text_len):
                extracted_text = response['Blocks'][max(text_len, key=text_len.get)]['Text']
                if extracted_text == '-':
                    extracted_text = ''
                else:
                    pass
            else:
                extracted_text = ''
            s3.Object(bucket_name,file_name).delete()
            return extracted_text
        except:
            return f"HTTP Status code is not 200 when running extract_text on {loc}, code is {status_code}"

优化方案

1. 跳过S3中转,直接传递图片字节

Textract的detect_document_text支持直接传入图片字节(Bytes参数),无需先上传到S3再删除,这能节省大量IO时间。修改提取函数去掉S3相关操作:

def extract_text(img, loc):
    try:
        # 转换图片为字节流
        img_obj = Image.fromarray(img).convert('RGB')
        out_img_obj = io.BytesIO()
        img_obj.save(out_img_obj, format="png")
        out_img_obj.seek(0)
        img_bytes = out_img_obj.getvalue()

        # 直接调用Textract,传入字节流
        response = textract_client.detect_document_text(
            Document={'Bytes': img_bytes}
        )

        # 提取最长文本(保留原逻辑)
        text_len = {}
        for block in response['Blocks']:
            if 'Text' in block:
                text_len[block['Id']] = len(block['Text'])
        
        if text_len:
            extracted_text = response['Blocks'][max(text_len, key=text_len.get)]['Text']
            return '' if extracted_text == '-' else extracted_text
        return ''
    except Exception as e:
        return f"Error extracting text for {loc}: {str(e)}"

2. 用多线程并行处理

由于调用Textract和图片转换属于IO密集型操作,使用多线程可以大幅提升效率。用concurrent.futures.ThreadPoolExecutor实现并行,同时打包每个任务的上下文(section、key、img、loc),确保结果能正确映射回原字典:

from concurrent.futures import ThreadPoolExecutor

def process_all_images(imgs):
    # 整理所有任务:每个任务包含section、key、img、loc
    tasks = []
    for section, items in imgs.items():
        for key, img in items.items():
            loc = f"{section.lower().replace(' ', '')}{key}"  # 生成类似sa1的loc
            tasks.append( (section, key, img, loc) )
    
    # 并行执行,设置线程数(可根据AWS配额调整,建议20-50)
    aws_output = {section: {} for section in imgs.keys()}
    with ThreadPoolExecutor(max_workers=30) as executor:
        # 提交所有任务,保留future与上下文的映射
        future_to_task = {executor.submit(extract_text, img, loc): (section, key) for section, key, img, loc in tasks}
        
        # 逐个获取结果并填充字典
        for future in future_to_task:
            section, key = future_to_task[future]
            try:
                result = future.result()
                aws_output[section][key] = result
            except Exception as e:
                aws_output[section][key] = f"Task failed: {str(e)}"
    
    return aws_output

# 使用示例
aws_output = process_all_images(imgs)

3. 注意事项

  • AWS Textract有调用配额,默认同步调用是1000次/秒,线程数不要超过配额限制,避免触发限流。如果需要更高并发,可以申请提升配额。
  • 若图片数量极大(数千张),可以考虑使用Textract的异步批量处理(StartDocumentTextDetection),但需要处理任务队列、轮询结果,复杂度稍高,适合超大规模场景。
  • 原代码中status_code == 200是无效判断,改为直接捕获异常即可,因为调用失败时boto3会直接抛出异常。

内容的提问来源于stack exchange,提问作者carousallie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 13:43:23