调用GPT-4o API同时传入截图与PDF,如何让模型参考PDF内容分析截图?
调用GPT-4o API同时传入截图与PDF,如何让模型参考PDF内容分析截图?
我明白你遇到的问题了——虽然能成功发送包含截图和PDF的请求,但模型完全没用到PDF的内容对吧?这大概率是因为你的API请求结构有问题,还有prompt的引导不够明确。下面我来帮你修正问题:
问题1:API请求的Payload结构错误
你在messages的单个user消息对象里加了files字段,但这个字段根本不属于Chat Completions API的消息结构!GPT-4o要同时处理图片和PDF,必须把PDF作为content数组里的一个独立元素(和图片、文本同级),而不是放在消息的额外字段里。
问题2:Prompt引导不够明确
你的prompt只是说“用附件PDF帮助分析截图”,但没有明确要求模型先读取并利用PDF的内容,模型可能会默认只处理图片。
修正后的完整代码
我已经帮你调整了代码结构,同时优化了prompt引导:
import base64 import os import requests import json # 补上你可能有的encode_image函数(假设你之前实现过但没贴出来) def encode_image(image_path): """Encodes image to base64""" try: with open(image_path, "rb") as image_file: return base64.b64encode(image_file.read()).decode('utf-8') except FileNotFoundError: print(f"Error: The file '{image_path}' was not found.") return None def encode_pdf_to_base64(pdf_path): """Encodes PDF to base64""" try: with open(pdf_path, "rb") as pdf_file: return base64.b64encode(pdf_file.read()).decode('utf-8') except FileNotFoundError: print(f"Error: The file '{pdf_path}' was not found.") return None def analyze_with_pdf_and_screenshot(screenshot_path, pdf_path=None, api_key=None): """Analyzes the Screenshot with the help of the PDF""" base64_image = encode_image(screenshot_path) if not base64_image: return "Error: Failed to encode screenshot" base64_pdf = encode_pdf_to_base64(pdf_path) if pdf_path else None # 构建正确的API请求Payload结构 payload = { "model": "gpt-4o", "messages": [ { "role": "user", "content": [ # 1. 明确提示模型先读取PDF再分析 { "type": "text", "text": """请先完整读取并理解附件PDF中的所有内容,然后结合下面的截图进行分析: 1. 严格参考PDF中的规则、数据或说明内容 2. 所有分析结论必须同时结合PDF和截图的信息,不得编造内容""" }, # 2. 嵌入PDF文件(仅当PDF路径存在时添加) *( [ { "type": "file", "file": { "name": os.path.basename(pdf_path), "type": "application/pdf", "content": base64_pdf } } ] if base64_pdf else [] ), # 3. 嵌入截图 { "type": "image_url", "image_url": { "url": f"data:image/jpeg;base64,{base64_image}" } } ] } ], "max_tokens": 2500 } try: headers = { "Content-Type": "application/json", "Authorization": f"Bearer {api_key}" } response = requests.post( "https://api.openai.com/v1/chat/completions", headers=headers, json=payload ) response.raise_for_status() result = response.json() print("API-Answer:", json.dumps(result, indent=2)) # 返回模型的实际回答内容 return result["choices"][0]["message"]["content"] except requests.RequestException as e: print(f"Error analyzing file: {e}") # 打印API错误详情,方便排查问题 if hasattr(e, 'response') and e.response is not None: print("API Error Response:", e.response.text) return ""
关键注意事项
- 明确的指令引导:一定要在文本提示里清晰告诉模型“先读取PDF内容,再结合截图分析”,GPT-4o不会自动优先处理PDF,必须给出明确指令。
- 合规的Content结构:PDF和截图都要作为
content数组里的独立元素,每个元素对应正确的type(file/image_url/text)。 - 错误处理优化:新增了对图片编码失败的判断,还增加了打印API错误响应的逻辑,方便你快速排查后续可能出现的问题。
- 权限与兼容性:确保你的API密钥有权限调用GPT-4o模型,且使用的OpenAI SDK/API版本支持
file类型的content元素。
你可以测试一下这个修正后的代码,应该就能让模型正确参考PDF内容来分析截图了!
备注:内容来源于stack exchange,提问作者Nico Morningstar
相关产品推荐
相关产品推荐

