Azure混合搜索调用OpenAI Token消耗过高的技术咨询
关于Azure OpenAI结合Azure Search的性能与Token消耗问题
我们在Microsoft Azure上使用OpenAI时,常规API调用的Token消耗为prompt_tokens 36、total_tokens 113;但调用Azure Search端点时,prompt_tokens飙升至6193、total_tokens达6225,Token消耗激增100倍。
常规API调用返回示例
{ choices: [ { content_filter_results: [Object], finish_reason: 'stop', index: 0, logprobs: null, message: [Object] } ], created: 1721653544, id: 'chatcmpl-9nn1UksGuyliAGQsnqUrh3ADv5UhR', model: 'gpt-4o-2024-05-13', object: 'chat.completion', prompt_filter_results: [ { prompt_index: 0, content_filter_results: [Object] } ], system_fingerprint: 'fp_abc28019ad', usage: { completion_tokens: 77, prompt_tokens: 36, total_tokens: 113 } }
搜索端点调用返回示例
{ id: 'afdcc11f-66e6-412b-8c03-5ee25a20d249', model: 'gpt-4o', created: 1721654963, object: 'extensions.chat.completion', choices: [ { index: 0, finish_reason: 'stop', message: [Object] } ], usage: { prompt_tokens: 6193, completion_tokens: 32, total_tokens: 6225 }, system_fingerprint: 'fp_abc28019ad' }
搜索API调用配置
const SEARCH_BODY_TEMPLATE = { data_sources: [ { type: "azure_search", parameters: { filter: null, endpoint: process.env.SEARCH_SERVICE_ENDPOINT, index_name: process.env.SEARCH_INDEX_NAME, project_resource_id: process.env.SEARCH_PROJECT_RESOURCE_ID, semantic_configuration: "azureml-default", authentication: { "type": "system_assigned_managed_identity", "key": null }, role_information: "Your name is POC. Your are an intelligent assistant that has been developed to help all employees. Keep your answers short and clear.", in_scope: true, strictness: 1, top_n_documents: 3, key: process.env.SEARCH_KEY, embedding_endpoint: "https://xxx.openai.azure.com/openai/deployments/text-embedding-ada-002/embeddings?api-version=2023-05-15", embedding_key: process.env.OPEN_AI_API_KEY, query_type: "vectorSimpleHybrid" } } ], messages: [{ role: "system", content: "You are a basic assitant. Answer only if you really know. Otherwise answer 'i don't know'." }, { role: "user", content: "What is math?" }], deployment: process.env.SEARCH_DEPLOYMENT_ID, temperature: 0.7, top_p: 0.95, max_tokens: 200, stop: null, frequency_penalty: 0, presence_penalty: 0, }
我们试过通过Azure Web UI拖拽PDF文件、调整上传文档数量、代码分块上传索引等优化方式,均未改善Token消耗过高的问题。
技术问题与解答
1. 如何提升性能以支持数百用户同时使用?
- 优化检索结果数量:当前配置
top_n_documents:3,可尝试降低到1或2,减少传入大模型的上下文内容,降低Token消耗同时提升响应速度。 - 精细化文档分块:调整分块大小至500-1000Token,并添加100Token左右的重叠窗口,确保语义完整性的同时减少单块内容长度。
- 缓存高频查询结果:针对用户常问问题,提前生成并缓存回答,避免重复调用搜索和大模型,降低资源消耗。
- 扩容Azure资源:将Azure Search服务升级到更高层级(如Standard),为OpenAI部署增加实例数或选择高吞吐量SKU,提升并发支撑能力。
- 异步请求处理:引入请求队列机制,避免瞬间请求压垮服务,同时使用异步调用让前端无需等待长响应。
2. 该Token消耗水平是否属于Azure Search的正常情况?
这种Token消耗水平不属于常规合理范围。结合Azure Search的RAG调用,prompt_tokens通常应控制在1000-3000左右,6000+的消耗说明传入大模型的上下文存在大量冗余:
- 文档分块过大,返回的检索结果内容过长;
- 配置中
role_information和system提示词重复定义助手角色,增加冗余Token; - 搜索返回的文档未做摘要处理,直接传入大模型。
3. Azure平台上是否有更高效的自有数据对话实现方案?
- Azure OpenAI内置RAG优化:使用
semantic_ranker和semantic_captions功能,让Search返回精准摘要而非完整文档,大幅减少上下文Token。 - 自定义RAG实现:不依赖OpenAI的extensions端点,自行实现检索-增强逻辑:先调用Azure Search获取文档,提取关键信息或生成摘要后再传入大模型,全程可控上下文内容。
- Azure AI Studio RAG模板:使用预构建的RAG项目模板,内置分块、检索优化、缓存等最佳实践,快速部署高效的自有数据对话系统。
内容的提问来源于stack exchange,提问作者p0rter
相关产品推荐
相关产品推荐

