本地Python脚本访问Google Colab远程Ollama模型的403错误解决方法
解决通过Python脚本访问远程Colab Ollama实例的403错误
问题背景
已在Google Colab上通过以下代码成功部署Ollama,本地终端通过设置export OLLAMA_HOST=https://{url}.ngrok-free.app/后,执行ollama run llama2可正常访问,但本地Python脚本发送POST请求时返回403错误。
Colab部署代码:
from google.colab import userdata NGROK_AUTH_TOKEN = userdata.get('NGROK_AUTH_TOKEN') # 下载并安装ollama !curl https://ollama.ai/install.sh | sh !pip install aiohttp pyngrok import os import asyncio # 设置NVIDIA库路径 os.environ.update({'LD_LIBRARY_PATH': '/usr/lib64-nvidia'}) async def run_process(cmd): print('>>> starting', *cmd) p = await asyncio.subprocess.create_subprocess_exec( *cmd, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE, ) async def pipe(lines): async for line in lines: print(line.strip().decode('utf-8')) await asyncio.gather( pipe(p.stdout), pipe(p.stderr), ) # 配置ngrok认证令牌 await asyncio.gather( run_process(['ngrok', 'config', 'add-authtoken', NGROK_AUTH_TOKEN]) ) # 启动Ollama服务和ngrok隧道 await asyncio.gather( run_process(['ollama', 'serve']), run_process(['ngrok', 'http', '--log', 'stderr', '11434']), )
错误的Python访问脚本:
import requests # Ngrok隧道URL ngrok_tunnel_url = "https://{url}.ngrok-free.app/" # 请求体 payload = { "model": "llama2", "prompt": "Why is the sky blue?" } try: # 发送POST请求 response = requests.post(ngrok_tunnel_url, json=payload) if response.status_code == 200: print("请求成功:") print("响应内容:") print(response.text) else: print("错误:", response.status_code) except requests.exceptions.RequestException as e: print("错误:", e)
错误原因
- API端点路径错误:Ollama的生成式API端点不是根路径,必须指定
/api/generate(单轮内容生成)或/api/chat(多轮对话),直接请求根URL会返回403。 - 流式响应处理缺失:Ollama默认以流式方式返回结果,普通的
requests.post不会自动处理流式数据,可能导致请求被拒绝或解析异常。
解决方法与正确代码示例
方法1:处理流式响应(推荐)
Ollama默认启用流式输出,需逐行读取响应内容:
import requests import json ngrok_tunnel_url = "https://{url}.ngrok-free.app/api/generate" payload = { "model": "llama2", "prompt": "为什么天空是蓝色的?", "stream": True # 显式开启流式,默认值为True可省略 } try: # 发送流式请求 with requests.post(ngrok_tunnel_url, json=payload, stream=True) as response: response.raise_for_status() # 抛出HTTP错误异常 print("响应内容:") for line in response.iter_lines(): if line: data = json.loads(line) print(data.get('response', ''), end='') if data.get('done'): break except requests.exceptions.RequestException as e: print("请求错误:", e)
方法2:禁用流式响应
若不需要实时输出,可在请求体中设置"stream": False,一次性获取完整结果:
import requests import json ngrok_tunnel_url = "https://{url}.ngrok-free.app/api/generate" payload = { "model": "llama2", "prompt": "为什么天空是蓝色的?", "stream": False } try: response = requests.post(ngrok_tunnel_url, json=payload) response.raise_for_status() data = response.json() print("完整响应:") print(data.get('response')) except requests.exceptions.RequestException as e: print("请求错误:", e) except json.JSONDecodeError: print("响应解析错误:", response.text)
额外注意事项
- 确保Colab中的Ollama服务和ngrok隧道处于运行状态,免费ngrok隧道24小时后会失效,需重新生成。
- 若仍出现403,检查ngrok控制台是否有请求被拦截,可尝试重启隧道服务。
内容的提问来源于stack exchange,提问作者s123
相关产品推荐
相关产品推荐

