You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调用Hugging Face API无响应,出现504网关超时错误求助

Hugging Face API 504网关超时问题排查与解决

问题背景

刚入门生成式AI,使用Hugging Face开源模型,已创建access token并配置到.env文件,但调用API时始终返回504网关超时错误,更换模型后问题依旧。

运行代码

from langchain_huggingface import ChatHuggingFace, HuggingFaceEndpoint
from dotenv import load_dotenv

load_dotenv()

llm = HuggingFaceEndpoint(
    repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0',
    task='text-generation'
)

model = ChatHuggingFace(llm=llm)

result = model.invoke("What is the capital of India?")

print(result.content)

核心报错信息

504 Server Error: Gateway Time-out for url: https://router.huggingface.co/featherless-ai/v1/chat/completions


解决方案

1. 直接指定模型推理端点

Hugging Face的自动路由服务器负载过高时容易超时,跳过自动路由,直接使用模型专属的推理端点:

llm = HuggingFaceEndpoint(
    repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0',
    task='text-generation',
    endpoint_url='https://api-inference.huggingface.co/models/TinyLlama/TinyLlama-1.1B-Chat-v1.0'
)

2. 延长超时时间

部分小模型首次调用需要加载,默认超时时间可能不够,手动设置更长的超时:

llm = HuggingFaceEndpoint(
    repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0',
    task='text-generation',
    timeout=300  # 设置为5分钟,可根据需求调整
)

3. 验证Token权限

检查Hugging Face个人设置中的Access Tokens页面,确保你的token拥有read权限(默认创建的token具备该权限,若被修改需重新生成)。

4. 排查网络问题

  • 用curl测试API连通性(替换YOUR_ACCESS_TOKEN为实际token):
curl -H "Authorization: Bearer YOUR_ACCESS_TOKEN" https://api-inference.huggingface.co/models/TinyLlama/TinyLlama-1.1B-Chat-v1.0 -d '{"inputs":"What is the capital of India?"}'
  • 若在受限网络环境下,添加代理配置:
import os
os.environ['HTTP_PROXY'] = 'http://你的代理地址:端口'
os.environ['HTTPS_PROXY'] = 'http://你的代理地址:端口'

5. 本地运行模型替代API调用

如果远程API持续不稳定,直接在本地加载模型运行,无需依赖远程服务:

from langchain_huggingface import ChatHuggingFace, HuggingFacePipeline
from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline

# 加载模型和tokenizer
tokenizer = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
model = AutoModelForCausalLM.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")

# 创建文本生成管道
pipe = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    max_new_tokens=100,
    temperature=0.7
)

# 包装为LangChain可调用对象
llm = HuggingFacePipeline(pipeline=pipe)
chat_model = ChatHuggingFace(llm=llm)

# 调用模型
result = chat_model.invoke("What is the capital of India?")
print(result.content)

内容的提问来源于stack exchange,提问作者Apoorva Walia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 09:13:18