如何在AWS Bedrock上使用Claude 3 API实现PDF文档摘要?
处理AWS Bedrock Claude 3 API中的PDF摘要问题
针对你提出的核心问题,直接给出明确解决方案:
1. 是否需要先提取PDF文本再发送给Claude?
是的,必须先提取PDF中的文本内容。当前Bedrock上的Claude 3系列模型(Sonnet/Opus/Haiku)不支持直接传入PDF文件作为输入,仅能接收文本或Base64编码的图片。你可以用Python的PyPDF2或pdfplumber库提取PDF文本,示例代码片段:
import pdfplumber def extract_pdf_text(pdf_path): text = "" with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: text += page.extract_text() or "" return text
2. 是否有类似图片的直接处理方式?
没有。Claude 3在Bedrock的API仅支持text和image类型的content输入,不支持直接上传或编码PDF文件作为输入项,因此必须先完成PDF文本提取,再将文本传入API。
3. PDF包含图片该如何处理?
分两种场景处理:
- 图片含文字(如扫描件、内嵌文字图片):先用OCR工具(比如
pytesseract)识别图片中的文字,将识别结果和PDF提取的文本合并后,一起发送给Claude生成摘要。 - 纯图形图片(无文字):将图片转为Base64编码,和PDF提取的文本一起放入请求的
content数组中,让Claude结合文本与图片内容生成摘要。
整合处理的示例代码
基于你现有的函数,扩展为支持带图片的PDF处理:
import json import pdfplumber import pytesseract from PIL import Image import base64 from io import BytesIO def extract_pdf_content(pdf_path): content = [] # 提取PDF文本 text = "" with pdfplumber.open(pdf_path) as pdf: for page in pdf.pages: # 提取页面文本 page_text = page.extract_text() or "" text += page_text # 处理页面中的图片 for img in page.images: img_obj = Image.open(BytesIO(img["stream"])) # 尝试OCR识别图片文字 img_text = pytesseract.image_to_string(img_obj) if img_text.strip(): text += "\n" + img_text else: # 纯图形图片转为Base64 buffer = BytesIO() img_obj.save(buffer, format="PNG") img_base64 = base64.b64encode(buffer.getvalue()).decode("utf-8") content.append({"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": img_base64}}) # 添加文本内容到最前面 content.insert(0, {"type": "text", "text": f"请对以下PDF内容生成摘要:\n{text}"}) return content def invoke_claude_3_with_pdf(self, pdf_path): client = self.bedrock_client model_id = "anthropic.claude-3-sonnet-20240229-v1:0" content = extract_pdf_content(pdf_path) try: response = client.invoke_model( modelId=model_id, body=json.dumps( { "anthropic_version": "bedrock-2023-05-31", "max_tokens": 2048, "messages": [ { "role": "user", "content": content, } ], } ), ) result = json.loads(response.get("body").read()) input_tokens = result["usage"]["input_tokens"] output_tokens = result["usage"]["output_tokens"] output_list = result.get("content", []) return output_list, input_tokens, output_tokens except Exception as e: print(f"API调用失败:{str(e)}") return None, 0, 0
内容的提问来源于stack exchange,提问作者Naxi
相关产品推荐
相关产品推荐

