本地运行Phi-3大模型遇HttpClient超时错误,寻求技术支持
本地Phi-3模型调用超时问题解决
我通过以下代码向本地LLM(Phi-3)发送提示请求,LM Studio中可正常返回响应,但在Visual Studio中收到错误:“请求因配置的HttpClient.Timeout(100秒)到期而被取消”,寻求解决协助。
代码如下:
var phi3 = new CustomChatCompletionService(); phi3.ModelUrl = "http://localhost:1234/v1/chat/completions"; // semantic kernel builder var builder = Kernel.CreateBuilder(); builder.Services.AddKeyedSingleton<IChatCompletionService>("microsoft/Phi-3-mini-4k-instruct-gguf", phi3); var kernel = builder.Build(); // init chat var chat = kernel.GetRequiredService<IChatCompletionService>(); var history = new ChatHistory(); history.AddSystemMessage("You are a useful assistant that replies using a funny style and emojis. Your name is Goku."); history.AddUserMessage("hi, who are you?"); // print response var result = await chat.GetChatMessageContentsAsync(history); Console.WriteLine(result[^1].Content);
解决建议:
- 确认本地服务状态:检查LM Studio的本地API服务(端口1234)是否正常运行,同时确保Visual Studio的网络请求未被防火墙或杀毒软件拦截。
- 延长HttpClient超时时间:在实例化CustomChatCompletionService时,显式配置更长的超时时长,避免因模型推理慢导致超时:
var phi3 = new CustomChatCompletionService(); phi3.ModelUrl = "http://localhost:1234/v1/chat/completions"; // 设置超时为5分钟,根据实际推理速度调整 phi3.HttpClient = new HttpClient { Timeout = TimeSpan.FromMinutes(5) }; - 优化模型推理性能:在LM Studio中调整模型参数,比如启用GPU硬件加速、降低上下文窗口大小、调整温度参数等,提升本地模型的响应速度。
- 验证API兼容性:用Postman或curl工具直接调用
http://localhost:1234/v1/chat/completions接口,发送与代码中相同的聊天请求,确认接口本身能正常快速返回,排查是否是代码层面的请求格式问题。 - 检查CustomChatCompletionService实现:确保自定义服务正确处理了请求序列化、头部设置等细节,与LM Studio的OpenAI兼容API格式匹配。
内容的提问来源于stack exchange,提问作者renakre
相关产品推荐
相关产品推荐

