如何在不依赖LLM决策的情况下从Semantic Kernel向MCP服务器传递动态参数(如索引名称、密钥)
如何在不依赖LLM决策的情况下从Semantic Kernel向MCP服务器传递动态参数(如索引名称、密钥)
针对你这个生产级Semantic Kernel + MCP的场景,我结合微软的企业实践给你整理了一套完全避开LLM决策的解决方案,涵盖上下文传递、插件修改、多连接管理等各个环节:
一、微软推荐的企业级上下文传递模式
在生产环境中,HTTP请求头+Kernel上下文拦截是最符合要求的方案,它完全绕开LLM参与路由决策,同时满足可扩展性、可审计、低延迟的需求:
- 确定性:索引选择逻辑在FastAPI层基于
client_id/bot_id直接映射(比如从配置中心或数据库拉取对应关系),100%可控 - 性能:仅在请求头添加少量字段,无额外性能开销
- 安全:可结合Azure AD认证验证请求合法性,请求头可添加签名防止篡改
- 可观测性:请求头里的
client_id、索引名等信息可直接接入Azure Monitor做审计和监控
二、修改MCPStreamableHttpPlugin实现动态参数传递
你可以通过扩展内置的MCPStreamableHttpPlugin,把动态上下文注入到请求头中,具体代码示例如下:
客户端(FastAPI + Semantic Kernel)修改
from semantic_kernel.connectors.mcp.mcp_streamable_http_plugin import MCPStreamableHttpPlugin from semantic_kernel.kernel_context import KernelContext # 扩展MCP插件,添加动态上下文注入逻辑 class ContextAwareMCPPlugin(MCPStreamableHttpPlugin): async def invoke(self, context: KernelContext) -> KernelFunctionResult: # 从KernelContext中取出提前存入的索引名、客户端ID index_name = context.variables.get("pinecone_index") client_id = context.variables.get("client_id") # 注入到HTTP请求头 self._http_client.headers.update({ "X-Pinecone-Index": index_name, "X-Client-ID": client_id }) return await super().invoke(context) # 在FastAPI的聊天接口中使用 async def chat_endpoint( request: ChatRequest, conversation_id: str = Depends(get_conversation_id), client_id: str = Depends(get_client_id), bot_id: str = Depends(get_bot_id) ): # 1. 根据client_id确定对应的索引名(可从配置中心/数据库拉取) pinecone_index = get_index_for_client(client_id) # 自定义映射逻辑 # 2. 将上下文存入KernelContext context = kernel.create_new_context() context.variables["pinecone_index"] = pinecone_index context.variables["client_id"] = client_id # 3. 使用自定义插件调用MCP服务 result_item = await agent.get_response( messages=request.message, thread=channel.thread, context=context )
MCP服务器(FastMCP)修改
from fastapi import Request @mcp.tool(structured_output=True) async def search_pinecone_tool(query: str, top_k: int = 5, request: Request = None) -> str: # 从请求头读取动态索引名(兜底默认值) index_name = request.headers.get("X-Pinecone-Index", "azure-docs-chat") client_id = request.headers.get("X-Client-ID") # 记录审计日志(对接企业监控工具) logger.info(f"Authorized client {client_id} accessing Pinecone index {index_name} | Query: {query}") # 使用缓存的索引客户端(避免重复初始化连接) pinecone_index = _get_cached_pinecone_index(index_name) # 执行向量搜索逻辑... search_results = pinecone_index.query(query=query, top_k=top_k) return str(search_results)
三、拦截MCP请求注入上下文的轻量化方案
如果不想重写插件,也可以用Semantic Kernel的KernelMiddleware拦截请求,在发送到MCP前注入上下文:
from semantic_kernel.middleware.kernel_middleware import KernelMiddleware from semantic_kernel.kernel_context import KernelContext class MCPContextInjectorMiddleware(KernelMiddleware): async def before_function_invocation(self, context: KernelContext, function) -> None: # 只拦截MCP插件的调用 if function.plugin_name == "MCPStreamableHttpPlugin": index_name = context.variables.get("pinecone_index") if index_name: # 注入请求头 function._http_client.headers["X-Pinecone-Index"] = index_name # 注册中间件到Kernel kernel.add_middleware(MCPContextInjectorMiddleware())
四、多MCP连接的最佳实践
针对多索引/多数据库的场景,推荐以下管理方式:
- 连接池缓存:为每个索引创建独立的Pinecone/Azure AI Search客户端实例,存入内存缓存(比如用
functools.lru_cache),避免每次请求都新建连接 - 配置中心化:把索引与客户端的映射关系、认证信息存在Azure App Configuration中,动态拉取,无需修改代码重启服务
- 租户隔离:多租户场景下,为每个租户分配独立的连接池,避免资源竞争和数据泄露
五、Semantic Kernel内置的上下文传递特性
你可能忽略了几个内置特性,可以直接用来传递上下文:
KernelContext:核心的上下文容器,可以存储任意键值对,在插件调用时自动传递FunctionInvocationContext:插件调用时可获取到完整的请求上下文,包括用户身份、会话ID等PluginMetadata:可以在定义MCP插件时添加自定义元数据,但不如请求头灵活直接
内容来源于stack exchange
相关产品推荐
相关产品推荐

