如何向Google Gemini发送并行请求批量提取图片文本?
批量并发处理图片文本提取方案
可以用Python的concurrent.futures.ThreadPoolExecutor实现并发请求,结合Gemini API每分钟60次的限制,设置最大并发数为60,同时处理多张图片,大幅缩短耗时。
修改后的代码实现
import os from PIL import Image import google.generativeai as genai from tqdm import tqdm from concurrent.futures import ThreadPoolExecutor, as_completed # 初始化Gemini模型 model = genai.GenerativeModel('gemini-pro-vision', safety_settings=safety_settings) image_dir = "你的图片目录路径" images_to_process = [os.path.join(image_dir, image_name) for image_name in os.listdir(image_dir)] prompt = """Carefully scan this image: if it has text, extract all the text and return the text from it. If the image does not have text return '<000>'.""" # 定义单张图片处理函数 def process_image(image_path): try: img = Image.open(image_path) output = model.generate_content([prompt, img]) output.resolve() # 确保请求完全完成 return image_path, output.text except Exception as e: return image_path, f"处理失败: {str(e)}" # 并发处理,设置最大并发数为60 results = [] with ThreadPoolExecutor(max_workers=60) as executor: # 提交所有处理任务 futures = {executor.submit(process_image, img_path): img_path for img_path in images_to_process} # 实时跟踪进度并收集结果 for future in tqdm(as_completed(futures), total=len(futures)): img_path, result = future.result() results.append((img_path, result)) print(f"{img_path}: {result}") # 可选:将提取结果保存到本地文件 with open("图片文本提取结果.txt", "w", encoding="utf-8") as f: for img_path, text in results: f.write(f"图片路径: {img_path}\n提取文本: {text}\n\n")
关键说明
- 并发控制:
max_workers=60确保同时最多发送60个请求,精准匹配API的每分钟请求限制,避免触发限流机制。 - 异常容错:为单张图片处理添加异常捕获,避免单个请求失败导致整个批量任务中断。
- 进度可视化:通过
as_completed配合tqdm,实时展示任务处理进度,直观了解整体推进情况。 - 请求完整性:调用
output.resolve()确保异步请求完全处理完毕,避免返回未完成的无效结果。
内容的提问来源于stack exchange,提问作者Adarsh Wase
相关产品推荐
相关产品推荐

