You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Langchain AgentExecutor调用Ollama Gemma模型报404错误的解决求助

解决Ollama+LangChain AgentExecutor调用时的404错误

问题背景

我按照ZenML的LLM Agents教程操作,做了以下修改:

  • 用Docker部署Ollama并运行gemma:2b模型
  • 将教程中的ChatOpenAI替换为该Ollama模型

修改steps/index_generator.py和steps/agent_creator.py后,成功运行Pipeline生成了langchain.agents.agent.AgentExecutor类型的Agent。但在另一个Python脚本中调用该Agent时:

from zenml import step, pipeline
from zenml.client import Client

client = Client()

agent = Client().get_artifact_version('86cb0da2-ca22-48ec-9548-410ccb073bc2') # 类型为langchain.agents.agent.AgentExecutor

question = "Hi!"

agent.run({"input": question,"chat_history": []})

出现错误:

OllamaEndpointNotFoundError: Ollama call failed with status code 404. Maybe your model is not found and you should 
pull the model with `ollama pull llama2`.

补充:我能通过CLI正常与gemma模型交互。

解决步骤

1. 检查Agent序列化时的Ollama配置

ZenML保存AgentExecutor artifact时,LangChain的Ollama模型配置(如模型名、服务地址)可能未被正确序列化:

  • 打开agent_creator.py,确保创建Ollama模型时明确指定model和base_url:
    from langchain.llms import Ollama
    
    llm = Ollama(
        model="gemma:2b",
        base_url="http://localhost:11434"  # 匹配Docker部署的Ollama地址,按需调整
    )
    
  • 重新运行Pipeline生成新的AgentExecutor artifact,再尝试调用。

2. 验证Ollama服务的网络可达性

如果调用脚本的环境无法访问Docker中的Ollama服务,会导致404错误:

  • 检查Docker容器的端口映射命令,确保已做端口暴露:
    docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
    
  • 在调用脚本所在环境执行curl http://localhost:11434/api/tags,确认能返回包含gemma:2b的模型列表;若为远程机器,将localhost替换为容器所在主机的IP。

3. 排查Agent中的模型名称残留

错误提示提到llama2,说明Agent可能仍在调用旧模型:

  • 检查agent_creator.py中所有与模型相关的配置,确保完全替换为gemma:2b,无ChatOpenAI或llama2的残留代码。
  • 创建Agent后可打印配置确认:
    print(agent.llm.model)  # 应输出gemma:2b
    

4. 绕过Artifact序列化问题重建Agent

若ZenML对AgentExecutor的序列化存在兼容性问题,可手动重建Agent:

from zenml import Client
from langchain.llms import Ollama
from langchain.agents import AgentExecutor

client = Client()
# 获取Agent的工具、prompt等配置 artifact
agent_config = client.get_artifact_version('your-agent-config-artifact-id')
# 重新初始化Ollama模型
llm = Ollama(model="gemma:2b", base_url="http://localhost:11434")
# 重建AgentExecutor
agent = AgentExecutor.from_agent_and_tools(
    agent=agent_config.agent,
    tools=agent_config.tools,
    verbose=True
)
# 调用测试
agent.run({"input": "Hi!", "chat_history": []})

内容的提问来源于stack exchange,提问作者happy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 00:52:37