使用Mistral API提取PDF表格遇阻:OCR无有效返回、Chat API无法访问URL
Mistral API提取PDF表格问题排查
问题描述
使用Mistral Le Chat功能可正常提取PDF表格,但通过API复现时遇到两个问题:
- OCR API调用返回结果为空,仅包含一张图片文件
- Chat Completion API返回提示:
I'm unable to directly access or view documents from URLs
原测试代码
第一段(OCR API调用)
from mistralai import Mistral from mistralai.models import File import os api_key = "API_KEY" client = Mistral(api_key=api_key) uploaded_pdf = await client.files.upload_async( file=File( file_name="table.pdf", content=open("./table.pdf", "rb").read(), ), purpose = "ocr", ) signed_url = client.files.get_signed_url(file_id=uploaded_pdf.id) ocr_response = client.ocr.process( model="mistral-ocr-latest", document={ "type": "document_url", "document_url": signed_url.url, } )
第二段(Chat Completion API调用)
from mistralai import Mistral from mistralai.models import File import os api_key = "API_KEY" client = Mistral(api_key=api_key) uploaded_pdf = await client.files.upload_async( file=File( file_name="table.pdf", content=open("./table.pdf", "rb").read(), ), purpose = "ocr", ) signed_url = client.files.get_signed_url(file_id=uploaded_pdf.id) # Define the messages for the chat messages = [ { "role": "user", "content": [ { "type": "text", "text": "Can you extract the data from this table from the PDF given." }, { "type": "document_url", "document_url": signed_url.url, } ] } ] # Get the chat response chat_response = client.chat.complete( model="mistral-large-latest", messages=messages, )
问题排查与解决方案
针对OCR API返回空结果的问题
- PDF类型适配:Mistral OCR API仅针对**扫描版PDF(图片型PDF)**生效,若你的PDF是原生可编辑PDF,使用OCR API会无法识别内容,建议改用文档处理流程。
- 异步方法使用规范:代码中使用
upload_async()异步上传方法,但如果运行环境是同步上下文,会导致上传未完成就执行后续操作,返回空结果。需改为同步方法upload(),或在异步函数中运行代码。 - 参数修正:确保上传文件时
purpose设置正确,扫描PDF用"ocr",原生PDF用"document-processing"。
修正后的OCR API调用代码(同步环境):
from mistralai import Mistral from mistralai.models import File import os api_key = "API_KEY" client = Mistral(api_key=api_key) # 同步上传文件 uploaded_pdf = client.files.upload( file=File( file_name="table.pdf", content=open("./table.pdf", "rb").read(), ), purpose = "ocr", # 仅扫描PDF使用该值 ) signed_url = client.files.get_signed_url(file_id=uploaded_pdf.id) ocr_response = client.ocr.process( model="mistral-ocr-latest", document={ "type": "document_url", "document_url": signed_url.url, } ) # 打印结果查看 print(ocr_response.text)
针对Chat Completion API无法访问URL的问题
- 文件引用方式错误:Mistral Chat Completion API不支持通过
document_url传递文件,需直接使用上传后的file_id引用文件。 - Purpose参数错误:上传文件时
purpose应设置为"document-processing",而非"ocr",否则模型无法读取文件内容。
修正后的Chat Completion API调用代码:
from mistralai import Mistral from mistralai.models import File import os api_key = "API_KEY" client = Mistral(api_key=api_key) # 上传文件,purpose设置为document-processing uploaded_pdf = client.files.upload( file=File( file_name="table.pdf", content=open("./table.pdf", "rb").read(), ), purpose = "document-processing", ) # 消息中直接使用file_id引用文件 messages = [ { "role": "user", "content": [ { "type": "text", "text": "提取这份PDF中的表格数据,以结构化格式返回。" }, { "type": "file", "file_id": uploaded_pdf.id } ] } ] chat_response = client.chat.complete( model="mistral-large-latest", messages=messages, ) # 打印提取结果 print(chat_response.choices[0].message.content)
内容的提问来源于stack exchange,提问作者Shelly Liu
相关产品推荐
相关产品推荐

