You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

本地运行Phi-3大模型遇HttpClient超时错误,寻求技术支持

本地Phi-3模型调用超时问题解决

我通过以下代码向本地LLM(Phi-3)发送提示请求,LM Studio中可正常返回响应,但在Visual Studio中收到错误:“请求因配置的HttpClient.Timeout(100秒)到期而被取消”,寻求解决协助。

代码如下:

var phi3 = new CustomChatCompletionService();
phi3.ModelUrl = "http://localhost:1234/v1/chat/completions";

// semantic kernel builder
var builder = Kernel.CreateBuilder();
builder.Services.AddKeyedSingleton<IChatCompletionService>("microsoft/Phi-3-mini-4k-instruct-gguf", phi3);
var kernel = builder.Build();

// init chat
var chat = kernel.GetRequiredService<IChatCompletionService>();
var history = new ChatHistory();
history.AddSystemMessage("You are a useful assistant that replies using a funny style and emojis. Your name is Goku.");
history.AddUserMessage("hi, who are you?");

// print response
var result = await chat.GetChatMessageContentsAsync(history);
Console.WriteLine(result[^1].Content);

解决建议:

  • 确认本地服务状态:检查LM Studio的本地API服务(端口1234)是否正常运行,同时确保Visual Studio的网络请求未被防火墙或杀毒软件拦截。
  • 延长HttpClient超时时间:在实例化CustomChatCompletionService时,显式配置更长的超时时长,避免因模型推理慢导致超时:
    var phi3 = new CustomChatCompletionService();
    phi3.ModelUrl = "http://localhost:1234/v1/chat/completions";
    // 设置超时为5分钟,根据实际推理速度调整
    phi3.HttpClient = new HttpClient { Timeout = TimeSpan.FromMinutes(5) };
    
  • 优化模型推理性能:在LM Studio中调整模型参数,比如启用GPU硬件加速、降低上下文窗口大小、调整温度参数等,提升本地模型的响应速度。
  • 验证API兼容性:用Postman或curl工具直接调用http://localhost:1234/v1/chat/completions接口,发送与代码中相同的聊天请求,确认接口本身能正常快速返回,排查是否是代码层面的请求格式问题。
  • 检查CustomChatCompletionService实现:确保自定义服务正确处理了请求序列化、头部设置等细节,与LM Studio的OpenAI兼容API格式匹配。

内容的提问来源于stack exchange,提问作者renakre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 21:46:16