调用Hugging Face API无响应,出现504网关超时错误求助
Hugging Face API 504网关超时问题排查与解决
问题背景
刚入门生成式AI,使用Hugging Face开源模型,已创建access token并配置到.env文件,但调用API时始终返回504网关超时错误,更换模型后问题依旧。
运行代码
from langchain_huggingface import ChatHuggingFace, HuggingFaceEndpoint from dotenv import load_dotenv load_dotenv() llm = HuggingFaceEndpoint( repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0', task='text-generation' ) model = ChatHuggingFace(llm=llm) result = model.invoke("What is the capital of India?") print(result.content)
核心报错信息
504 Server Error: Gateway Time-out for url: https://router.huggingface.co/featherless-ai/v1/chat/completions
解决方案
1. 直接指定模型推理端点
Hugging Face的自动路由服务器负载过高时容易超时,跳过自动路由,直接使用模型专属的推理端点:
llm = HuggingFaceEndpoint( repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0', task='text-generation', endpoint_url='https://api-inference.huggingface.co/models/TinyLlama/TinyLlama-1.1B-Chat-v1.0' )
2. 延长超时时间
部分小模型首次调用需要加载,默认超时时间可能不够,手动设置更长的超时:
llm = HuggingFaceEndpoint( repo_id='TinyLlama/TinyLlama-1.1B-Chat-v1.0', task='text-generation', timeout=300 # 设置为5分钟,可根据需求调整 )
3. 验证Token权限
检查Hugging Face个人设置中的Access Tokens页面,确保你的token拥有read权限(默认创建的token具备该权限,若被修改需重新生成)。
4. 排查网络问题
- 用curl测试API连通性(替换
YOUR_ACCESS_TOKEN为实际token):
curl -H "Authorization: Bearer YOUR_ACCESS_TOKEN" https://api-inference.huggingface.co/models/TinyLlama/TinyLlama-1.1B-Chat-v1.0 -d '{"inputs":"What is the capital of India?"}'
- 若在受限网络环境下,添加代理配置:
import os os.environ['HTTP_PROXY'] = 'http://你的代理地址:端口' os.environ['HTTPS_PROXY'] = 'http://你的代理地址:端口'
5. 本地运行模型替代API调用
如果远程API持续不稳定,直接在本地加载模型运行,无需依赖远程服务:
from langchain_huggingface import ChatHuggingFace, HuggingFacePipeline from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline # 加载模型和tokenizer tokenizer = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0") model = AutoModelForCausalLM.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0") # 创建文本生成管道 pipe = pipeline( "text-generation", model=model, tokenizer=tokenizer, max_new_tokens=100, temperature=0.7 ) # 包装为LangChain可调用对象 llm = HuggingFacePipeline(pipeline=pipe) chat_model = ChatHuggingFace(llm=llm) # 调用模型 result = chat_model.invoke("What is the capital of India?") print(result.content)
内容的提问来源于stack exchange,提问作者Apoorva Walia
相关产品推荐
相关产品推荐

