Azure OpenAI gpt-4o调用CompleteChatAsync失败排查求助
问题描述
6个月前正常运行的RAG Web服务,因SDK更新、模型迭代后彻底失效。调用gpt-4o模型时存在以下异常:
- 未使用
CancellationToken,CompleteChatAsync方法始终无返回; - 使用
CancellationToken,30秒后任务被取消,抛出“任务已取消”或“资源未找到”异常。
相同代码调用gpt-35-turbo-16k模型可正常运行。
环境与配置信息:
- 资源组及Azure OpenAI资源位于eastus区域,gpt-4o-2024-08-06模型已成功部署;
- 运行环境:Windows 11、Visual Studio 2022、.NET 9;
- 核心代码如下:
public async Task<string?> GenerateResponseAsync(CompletionModel model, string prompt, List<string> contextList, CancellationToken cancellationToken) { _logger.LogInformation("GenerateResponseAsync entered"); ArgumentNullException.ThrowIfNull(model, nameof(model)); ArgumentNullException.ThrowIfNullOrWhiteSpace(prompt, nameof(prompt)); string respMessage = ""; try { Uri azureOpenAIResourceUri = new Uri(_azureSettings.AzureOpenAiEndpoint); AzureKeyCredential azureOpenAiApiKey = new AzureKeyCredential(_azureSettings.AzureOpenAiApiKey); AzureOpenAIClient azureClient = new(azureOpenAIResourceUri, azureOpenAiApiKey); ChatClient chatClient = azureClient.GetChatClient(model.DeploymentName); if (chatClient == null) { throw new Exception("Failed to create AzureOpenAIClient"); } List<ChatMessage>? chatMessages = ConstructMessages(model, prompt, contextList); if (chatMessages == null) { throw new Exception("Failed to construct messages"); } _logger.LogInformation("Creating ChatCompletionOptions..."); ChatCompletionOptions? chatCompletionOptions = new ChatCompletionOptions() { FrequencyPenalty = 0f, PresencePenalty = 0f, MaxOutputTokenCount = model.TokenLimit - 1000, TopP = 0, Temperature = 1.0f, }; using var timeoutTokenSource = new CancellationTokenSource(TimeSpan.FromSeconds(30)); var linkedTokenSource = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken, timeoutTokenSource.Token); _logger.LogInformation("Calling CompleteChatAsync..."); ChatCompletion chatCompletion = await chatClient.CompleteChatAsync(chatMessages.ToArray(), chatCompletionOptions, linkedTokenSource.Token); if (chatCompletion != null) { ChatMessageContent msgContent = chatCompletion.Content; if (msgContent != null) { respMessage = (msgContent.Count > 0) ? msgContent[0].Text : ""; } } _logger.LogInformation($"CONTENT: {respMessage}"); } catch (Exception ex) { _logger.LogError(ex, ex.Message); throw; } finally { _logger.LogInformation("GenerateResponse exiting"); } return respMessage; } public List<ChatMessage>? ConstructMessages(CompletionModel model, string prompt, List<string> contextList) { _logger.LogInformation("ConstructMessages entered"); try { StringBuilder sb = new StringBuilder($"QUESTION: {prompt}"); sb.AppendLine(Environment.NewLine); sb.AppendLine("CONTEXT:"); sb.AppendLine(Environment.NewLine); if (contextList != null) { for (int i = 0; i < contextList.Count; i++) { sb.AppendLine("## " + contextList[i].Trim('\n').Trim('\r')); sb.AppendLine(Environment.NewLine); } } List<ChatMessage> chatMessages = new List<ChatMessage>() { new SystemChatMessage(_azureSettings.AzureOpenAiCompletionSystemInstruction), new UserChatMessage(sb.ToString()), }; _logger.LogInformation($"SystemInstruction: {chatMessages[0]}"); _logger.LogInformation($"User Message: {chatMessages[1]}"); string totPayloadStr = _azureSettings.AzureOpenAiCompletionSystemInstruction + sb.ToString(); if (IsExceedsModelTokenLimit(totPayloadStr, model) == true) { throw new Exception("Total prompt + system + context message size exceeds the model limit"); } _logger.LogInformation("Exiting exiting..."); return chatMessages; } catch(Exception ex) { _logger.LogError($"ConstructMessages exception: {ex.Message}"); } return null; }
排查与修复方案
1. 匹配SDK与.NET版本兼容性
.NET 9为较新框架,需确保使用的Azure.AI.OpenAI SDK版本与.NET 9兼容。建议升级至最新稳定版(如1.0.0-beta.17及以上),旧版SDK对gpt-4o模型的支持可能存在缺陷。
2. 调整超时配置
gpt-4o模型处理复杂请求的耗时通常长于gpt-35-turbo,30秒超时可能不足以完成响应:
- 延长超时时间至60秒或更久,测试是否能正常返回;
- 若无需额外超时控制,可直接使用传入的
cancellationToken,避免手动创建超时令牌导致提前取消。
修改示例:
// 延长超时时间至60秒 using var timeoutTokenSource = new CancellationTokenSource(TimeSpan.FromSeconds(60)); var linkedTokenSource = CancellationTokenSource.CreateLinkedTokenSource(cancellationToken, timeoutTokenSource.Token);
3. 验证模型部署与端点配置
- 确认gpt-4o-2024-08-06的部署名称与代码中
model.DeploymentName完全一致,注意大小写区分; - 检查Azure OpenAI端点格式是否正确,需符合
https://<resource-name>.openai.azure.com/规范; - 在Azure门户中测试该模型部署的在线推理,确认服务本身无故障。
4. 修正请求参数合法性
gpt-4o对部分参数的取值范围有明确要求,需调整非法配置:
TopP设置为0不符合规范,需调整为0.1-1.0之间的有效值;- 确认
MaxOutputTokenCount未超过gpt-4o模型的最大输出令牌限制(gpt-4o-2024-08-06默认最大输出为4096,部分部署支持16384,需匹配实际配置)。
修改示例:
ChatCompletionOptions? chatCompletionOptions = new ChatCompletionOptions() { FrequencyPenalty = 0f, PresencePenalty = 0f, MaxOutputTokenCount = Math.Min(model.TokenLimit - 1000, 4096), // 匹配模型实际输出限制 TopP = 0.9f, // 设置合法取值 Temperature = 1.0f, };
5. 优化令牌计算逻辑
- 确保
IsExceedsModelTokenLimit方法使用gpt-4o的令牌编码规则(不同模型的令牌计算方式存在差异),避免误判或实际超出上下文窗口导致服务无响应; - 检查系统提示词长度,过度冗长的提示词会增加处理耗时,建议精简非必要内容。
内容的提问来源于stack exchange,提问作者securigy
相关产品推荐
相关产品推荐

