如何让LLM通过Python调用Google Cloud Run上的MCP工具函数
实现LLM直接调用Cloud Run上的FastMCP加法服务
1. 配置Cloud Run访问权限
- 创建Google Cloud服务账号,赋予
roles/run.invoker角色(遵循最小权限原则,仅允许调用目标Cloud Run服务) - 为该服务账号生成JSON格式密钥文件,保存至安全位置(后续通过环境变量加载,禁止硬编码到代码中)
2. 确认Cloud Run服务部署配置
- 确保FastMCP容器暴露的端口与Cloud Run部署时指定的端口一致(例如FastMCP默认用8000端口,部署时需设置
--port 8000) - 验证服务可用性:通过
gcloud run services describe <service-name>获取服务URL,发送测试POST请求确认加法功能正常响应
3. 定义LLM的工具调用Schema
给GPT-4o等LLM明确工具描述,让它能自主判断何时调用加法服务,示例Schema如下:
{ "type": "function", "function": { "name": "cloud_addition_service", "description": "调用部署在Cloud Run上的FastMCP服务计算两个数字的和", "parameters": { "type": "object", "properties": { "a": { "type": "number", "description": "第一个加数" }, "b": { "type": "number", "description": "第二个加数" } }, "required": ["a", "b"] } } }
4. 编写工具调用逻辑(Python示例)
4.1 生成Cloud Run访问Token
基于JSON密钥生成OAuth2 Bearer Token,用于Cloud Run身份验证:
import google.auth from google.auth.transport.requests import Request def get_cloud_run_token(): credentials, _ = google.auth.load_credentials_from_file("service-account-key.json") credentials.refresh(Request()) return credentials.token
4.2 调用Cloud Run上的FastMCP服务
当LLM返回工具调用指令时,发送POST请求到服务URL:
import requests def call_cloud_addition(a, b, cloud_run_url): token = get_cloud_run_token() headers = { "Authorization": f"Bearer {token}", "Content-Type": "application/json" } payload = {"a": a, "b": b} response = requests.post(cloud_run_url, json=payload, headers=headers) response.raise_for_status() return response.json()["result"] # 假设FastMCP返回格式为{"result": 计算结果}
4.3 整合LLM对话流程
在LLM交互循环中,检测工具调用请求,执行调用后将结果返回给LLM继续生成响应:
from openai import OpenAI import os client = OpenAI() def chat_with_llm(user_query): tools = [ { "type": "function", "function": { "name": "cloud_addition_service", "description": "调用Cloud Run上的FastMCP服务计算两数之和", "parameters": { "type": "object", "properties": { "a": {"type": "number", "description": "第一个加数"}, "b": {"type": "number", "description": "第二个加数"} }, "required": ["a", "b"] } } } ] response = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": user_query}], tools=tools, tool_choice="auto" ) # 处理工具调用请求 if response.choices[0].finish_reason == "tool_calls": tool_call = response.choices[0].message.tool_calls[0] if tool_call.function.name == "cloud_addition_service": args = eval(tool_call.function.arguments) result = call_cloud_addition(args["a"], args["b"], os.environ.get("CLOUD_RUN_URL")) # 将工具结果返回给LLM生成最终响应 follow_up_response = client.chat.completions.create( model="gpt-4o", messages=[ {"role": "user", "content": user_query}, response.choices[0].message, { "role": "tool", "tool_call_id": tool_call.id, "content": str(result) } ] ) return follow_up_response.choices[0].message.content else: return response.choices[0].message.content
5. 安全与优化建议
- 通过环境变量加载服务账号密钥:
google.auth.load_credentials_from_file(os.environ.get("GOOGLE_APPLICATION_CREDENTIALS")) - 启用Cloud Run VPC连接器,限制服务仅允许内部或指定IP访问
- 定期轮换服务账号密钥,避免泄露风险
- 配置Cloud Run请求限流和配额,防止恶意调用
内容的提问来源于stack exchange,提问作者Sachu
相关产品推荐
相关产品推荐

