You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不依赖LLM决策的情况下从Semantic Kernel向MCP服务器传递动态参数(如索引名称、密钥)

如何在不依赖LLM决策的情况下从Semantic Kernel向MCP服务器传递动态参数(如索引名称、密钥)

针对你这个生产级Semantic Kernel + MCP的场景,我结合微软的企业实践给你整理了一套完全避开LLM决策的解决方案,涵盖上下文传递、插件修改、多连接管理等各个环节:


一、微软推荐的企业级上下文传递模式

在生产环境中,HTTP请求头+Kernel上下文拦截是最符合要求的方案,它完全绕开LLM参与路由决策,同时满足可扩展性、可审计、低延迟的需求:

  • 确定性:索引选择逻辑在FastAPI层基于client_id/bot_id直接映射(比如从配置中心或数据库拉取对应关系),100%可控
  • 性能:仅在请求头添加少量字段,无额外性能开销
  • 安全:可结合Azure AD认证验证请求合法性,请求头可添加签名防止篡改
  • 可观测性:请求头里的client_id、索引名等信息可直接接入Azure Monitor做审计和监控

二、修改MCPStreamableHttpPlugin实现动态参数传递

你可以通过扩展内置的MCPStreamableHttpPlugin,把动态上下文注入到请求头中,具体代码示例如下:

客户端(FastAPI + Semantic Kernel)修改

from semantic_kernel.connectors.mcp.mcp_streamable_http_plugin import MCPStreamableHttpPlugin
from semantic_kernel.kernel_context import KernelContext

# 扩展MCP插件,添加动态上下文注入逻辑
class ContextAwareMCPPlugin(MCPStreamableHttpPlugin):
    async def invoke(self, context: KernelContext) -> KernelFunctionResult:
        # 从KernelContext中取出提前存入的索引名、客户端ID
        index_name = context.variables.get("pinecone_index")
        client_id = context.variables.get("client_id")
        
        # 注入到HTTP请求头
        self._http_client.headers.update({
            "X-Pinecone-Index": index_name,
            "X-Client-ID": client_id
        })
        
        return await super().invoke(context)

# 在FastAPI的聊天接口中使用
async def chat_endpoint(
    request: ChatRequest,
    conversation_id: str = Depends(get_conversation_id),
    client_id: str = Depends(get_client_id),
    bot_id: str = Depends(get_bot_id)
):
    # 1. 根据client_id确定对应的索引名(可从配置中心/数据库拉取)
    pinecone_index = get_index_for_client(client_id)  # 自定义映射逻辑
    
    # 2. 将上下文存入KernelContext
    context = kernel.create_new_context()
    context.variables["pinecone_index"] = pinecone_index
    context.variables["client_id"] = client_id
    
    # 3. 使用自定义插件调用MCP服务
    result_item = await agent.get_response(
        messages=request.message,
        thread=channel.thread,
        context=context
    )

MCP服务器(FastMCP)修改

from fastapi import Request

@mcp.tool(structured_output=True)
async def search_pinecone_tool(query: str, top_k: int = 5, request: Request = None) -> str:
    # 从请求头读取动态索引名(兜底默认值)
    index_name = request.headers.get("X-Pinecone-Index", "azure-docs-chat")
    client_id = request.headers.get("X-Client-ID")
    
    # 记录审计日志(对接企业监控工具)
    logger.info(f"Authorized client {client_id} accessing Pinecone index {index_name} | Query: {query}")
    
    # 使用缓存的索引客户端(避免重复初始化连接)
    pinecone_index = _get_cached_pinecone_index(index_name)
    
    # 执行向量搜索逻辑...
    search_results = pinecone_index.query(query=query, top_k=top_k)
    return str(search_results)

三、拦截MCP请求注入上下文的轻量化方案

如果不想重写插件,也可以用Semantic Kernel的KernelMiddleware拦截请求,在发送到MCP前注入上下文:

from semantic_kernel.middleware.kernel_middleware import KernelMiddleware
from semantic_kernel.kernel_context import KernelContext

class MCPContextInjectorMiddleware(KernelMiddleware):
    async def before_function_invocation(self, context: KernelContext, function) -> None:
        # 只拦截MCP插件的调用
        if function.plugin_name == "MCPStreamableHttpPlugin":
            index_name = context.variables.get("pinecone_index")
            if index_name:
                # 注入请求头
                function._http_client.headers["X-Pinecone-Index"] = index_name

# 注册中间件到Kernel
kernel.add_middleware(MCPContextInjectorMiddleware())

四、多MCP连接的最佳实践

针对多索引/多数据库的场景,推荐以下管理方式:

  • 连接池缓存:为每个索引创建独立的Pinecone/Azure AI Search客户端实例,存入内存缓存(比如用functools.lru_cache),避免每次请求都新建连接
  • 配置中心化:把索引与客户端的映射关系、认证信息存在Azure App Configuration中,动态拉取,无需修改代码重启服务
  • 租户隔离:多租户场景下,为每个租户分配独立的连接池,避免资源竞争和数据泄露

五、Semantic Kernel内置的上下文传递特性

你可能忽略了几个内置特性,可以直接用来传递上下文:

  • KernelContext:核心的上下文容器,可以存储任意键值对,在插件调用时自动传递
  • FunctionInvocationContext:插件调用时可获取到完整的请求上下文,包括用户身份、会话ID等
  • PluginMetadata:可以在定义MCP插件时添加自定义元数据,但不如请求头灵活直接

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 09:24:54