You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Langchain调用HuggingFace的google/flan-t5-xl模型超时错误求助

解决Langchain调用HuggingFace模型超时问题

你运行的代码如下:

# https://github.com/hwchase17/langchain/blob/0e763677e4c334af80f2b542cb269f3786d8403f/docs/modules/models/llms/integrations/huggingface_hub.ipynb

from langchain import HuggingFaceHub, LLMChain
import os

hugging_face_write = "MY_KEY"
os.environ['HUGGINGFACEHUB_API_TOKEN'] = hugging_face_write

from langchain import PromptTemplate, HuggingFaceHub, LLMChain

template = """Question: {question}

Answer: Let's think step by step."""
prompt = PromptTemplate(template=template, input_variables=["question"])
llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub(repo_id="google/flan-t5-xl", model_kwargs={"temperature":0, "max_length":64}))

question = "What NFL team won the Super Bowl in the year Justin Beiber was born?"

print(llm_chain.run(question))

运行后出现的错误信息:

ValueError                                Traceback (most recent call last)
g:\Meine Ablage\python\lang_chain\langchain_huggingface_example.py in line 1
----> 19 print(llm_chain.run(question))

File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:213, in Chain.run(self, *args, **kwargs)
    211     if len(args) != 1:
    212         raise ValueError("`run` supports only one positional argument.")
--> 213     return self(args[0])[self.output_keys[0]]
    215 if kwargs and not args:
    216     return self(kwargs)[self.output_keys[0]]

File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:116, in Chain.__call__(self, inputs, return_only_outputs)
    114 except (KeyboardInterrupt, Exception) as e:
    115     self.callback_manager.on_chain_error(e, verbose=self.verbose)
--> 116     raise e
    117 self.callback_manager.on_chain_end(outputs, verbose=self.verbose)
    118 return self.prep_outputs(inputs, outputs, return_only_outputs)

File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:113, in Chain.__call__(self, inputs, return_only_outputs)
    107 self.callback_manager.on_chain_start(
    108     {"name": self.__class__.__name__},
    109     inputs,
    110     verbose=self.verbose,
    111 )
...
    106 if self.client.task == "text-generation":
    107     # Text generation return includes the starter text.
    108     text = response[0]["generated_text"][len(prompt) :]

ValueError: Error raised by inference API: Model google/flan-t5-xl time out

问题原因及解决方法

  • 模型负载过高:google/flan-t5-xl是热门大模型,HuggingFace免费推理API资源有限,高峰期容易出现超时。
  • 方案1:延长超时等待时长
    在初始化HuggingFaceHub时添加timeout参数,比如设置为300秒:
    llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub(
        repo_id="google/flan-t5-xl", 
        model_kwargs={"temperature":0, "max_length":64},
        timeout=300
    ))
    
  • 方案2:切换轻量模型
    改用google/flan-t5-small或google/flan-t5-base这类小模型,推理速度更快,不易超时:
    llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub(
        repo_id="google/flan-t5-small", 
        model_kwargs={"temperature":0, "max_length":64}
    ))
    
  • 方案3:本地部署模型
    若需频繁测试,建议用本地加载模型的方式,避免依赖远程API:
    from langchain.llms import HuggingFacePipeline
    from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline
    
    tokenizer = AutoTokenizer.from_pretrained("google/flan-t5-xl")
    model = AutoModelForSeq2SeqLM.from_pretrained("google/flan-t5-xl")
    pipe = pipeline(
        "text2text-generation",
        model=model,
        tokenizer=tokenizer,
        max_length=64,
        temperature=0
    )
    llm = HuggingFacePipeline(pipeline=pipe)
    llm_chain = LLMChain(prompt=prompt, llm=llm)
    print(llm_chain.run(question))
    
  • 方案4:检查网络稳定性
    确保网络能稳定访问HuggingFace推理服务,必要时切换网络环境。

内容的提问来源于stack exchange,提问作者Jennie94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:02:12