You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用GPT-4o API同时传入截图与PDF,如何让模型参考PDF内容分析截图?

调用GPT-4o API同时传入截图与PDF,如何让模型参考PDF内容分析截图?

我明白你遇到的问题了——虽然能成功发送包含截图和PDF的请求,但模型完全没用到PDF的内容对吧?这大概率是因为你的API请求结构有问题,还有prompt的引导不够明确。下面我来帮你修正问题:

问题1:API请求的Payload结构错误

你在messages的单个user消息对象里加了files字段,但这个字段根本不属于Chat Completions API的消息结构!GPT-4o要同时处理图片和PDF,必须把PDF作为content数组里的一个独立元素(和图片、文本同级),而不是放在消息的额外字段里。

问题2:Prompt引导不够明确

你的prompt只是说“用附件PDF帮助分析截图”,但没有明确要求模型先读取并利用PDF的内容,模型可能会默认只处理图片。

修正后的完整代码

我已经帮你调整了代码结构,同时优化了prompt引导:

import base64
import os
import requests
import json

# 补上你可能有的encode_image函数(假设你之前实现过但没贴出来)
def encode_image(image_path):
    """Encodes image to base64"""
    try:
        with open(image_path, "rb") as image_file:
            return base64.b64encode(image_file.read()).decode('utf-8')
    except FileNotFoundError:
        print(f"Error: The file '{image_path}' was not found.")
        return None

def encode_pdf_to_base64(pdf_path):
    """Encodes PDF to base64"""
    try:
        with open(pdf_path, "rb") as pdf_file:
            return base64.b64encode(pdf_file.read()).decode('utf-8')
    except FileNotFoundError:
        print(f"Error: The file '{pdf_path}' was not found.")
        return None

def analyze_with_pdf_and_screenshot(screenshot_path, pdf_path=None, api_key=None):
    """Analyzes the Screenshot with the help of the PDF"""
    base64_image = encode_image(screenshot_path)
    if not base64_image:
        return "Error: Failed to encode screenshot"
    
    base64_pdf = encode_pdf_to_base64(pdf_path) if pdf_path else None

    # 构建正确的API请求Payload结构
    payload = {
        "model": "gpt-4o",
        "messages": [
            {
                "role": "user",
                "content": [
                    # 1. 明确提示模型先读取PDF再分析
                    {
                        "type": "text",
                        "text": """请先完整读取并理解附件PDF中的所有内容,然后结合下面的截图进行分析:
                        1. 严格参考PDF中的规则、数据或说明内容
                        2. 所有分析结论必须同时结合PDF和截图的信息,不得编造内容"""
                    },
                    # 2. 嵌入PDF文件(仅当PDF路径存在时添加)
                    *(
                        [
                            {
                                "type": "file",
                                "file": {
                                    "name": os.path.basename(pdf_path),
                                    "type": "application/pdf",
                                    "content": base64_pdf
                                }
                            }
                        ] if base64_pdf else []
                    ),
                    # 3. 嵌入截图
                    {
                        "type": "image_url",
                        "image_url": {
                            "url": f"data:image/jpeg;base64,{base64_image}"
                        }
                    }
                ]
            }
        ],
        "max_tokens": 2500
    }

    try:
        headers = {
            "Content-Type": "application/json",
            "Authorization": f"Bearer {api_key}"
        }
        response = requests.post(
            "https://api.openai.com/v1/chat/completions",
            headers=headers,
            json=payload
        )
        response.raise_for_status()
        result = response.json()
        print("API-Answer:", json.dumps(result, indent=2))
        # 返回模型的实际回答内容
        return result["choices"][0]["message"]["content"]

    except requests.RequestException as e:
        print(f"Error analyzing file: {e}")
        # 打印API错误详情,方便排查问题
        if hasattr(e, 'response') and e.response is not None:
            print("API Error Response:", e.response.text)
        return ""

关键注意事项

  • 明确的指令引导:一定要在文本提示里清晰告诉模型“先读取PDF内容,再结合截图分析”,GPT-4o不会自动优先处理PDF,必须给出明确指令。
  • 合规的Content结构:PDF和截图都要作为content数组里的独立元素,每个元素对应正确的type(file/image_url/text)。
  • 错误处理优化:新增了对图片编码失败的判断,还增加了打印API错误响应的逻辑,方便你快速排查后续可能出现的问题。
  • 权限与兼容性:确保你的API密钥有权限调用GPT-4o模型,且使用的OpenAI SDK/API版本支持file类型的content元素。

你可以测试一下这个修正后的代码,应该就能让模型正确参考PDF内容来分析截图了!

备注:内容来源于stack exchange,提问作者Nico Morningstar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 16:34:30