使用Langchain调用HuggingFace的google/flan-t5-xl模型超时错误求助
解决Langchain调用HuggingFace模型超时问题
你运行的代码如下:
# https://github.com/hwchase17/langchain/blob/0e763677e4c334af80f2b542cb269f3786d8403f/docs/modules/models/llms/integrations/huggingface_hub.ipynb from langchain import HuggingFaceHub, LLMChain import os hugging_face_write = "MY_KEY" os.environ['HUGGINGFACEHUB_API_TOKEN'] = hugging_face_write from langchain import PromptTemplate, HuggingFaceHub, LLMChain template = """Question: {question} Answer: Let's think step by step.""" prompt = PromptTemplate(template=template, input_variables=["question"]) llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub(repo_id="google/flan-t5-xl", model_kwargs={"temperature":0, "max_length":64})) question = "What NFL team won the Super Bowl in the year Justin Beiber was born?" print(llm_chain.run(question))
运行后出现的错误信息:
ValueError Traceback (most recent call last) g:\Meine Ablage\python\lang_chain\langchain_huggingface_example.py in line 1 ----> 19 print(llm_chain.run(question)) File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:213, in Chain.run(self, *args, **kwargs) 211 if len(args) != 1: 212 raise ValueError("`run` supports only one positional argument.") --> 213 return self(args[0])[self.output_keys[0]] 215 if kwargs and not args: 216 return self(kwargs)[self.output_keys[0]] File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:116, in Chain.__call__(self, inputs, return_only_outputs) 114 except (KeyboardInterrupt, Exception) as e: 115 self.callback_manager.on_chain_error(e, verbose=self.verbose) --> 116 raise e 117 self.callback_manager.on_chain_end(outputs, verbose=self.verbose) 118 return self.prep_outputs(inputs, outputs, return_only_outputs) File c:\Users\johan\.conda\envs\lang_chain\Lib\site-packages\langchain\chains\base.py:113, in Chain.__call__(self, inputs, return_only_outputs) 107 self.callback_manager.on_chain_start( 108 {"name": self.__class__.__name__}, 109 inputs, 110 verbose=self.verbose, 111 ) ... 106 if self.client.task == "text-generation": 107 # Text generation return includes the starter text. 108 text = response[0]["generated_text"][len(prompt) :] ValueError: Error raised by inference API: Model google/flan-t5-xl time out
问题原因及解决方法
- 模型负载过高:google/flan-t5-xl是热门大模型,HuggingFace免费推理API资源有限,高峰期容易出现超时。
- 方案1:延长超时等待时长
在初始化HuggingFaceHub时添加timeout参数,比如设置为300秒:llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub( repo_id="google/flan-t5-xl", model_kwargs={"temperature":0, "max_length":64}, timeout=300 )) - 方案2:切换轻量模型
改用google/flan-t5-small或google/flan-t5-base这类小模型,推理速度更快,不易超时:llm_chain = LLMChain(prompt=prompt, llm=HuggingFaceHub( repo_id="google/flan-t5-small", model_kwargs={"temperature":0, "max_length":64} )) - 方案3:本地部署模型
若需频繁测试,建议用本地加载模型的方式,避免依赖远程API:from langchain.llms import HuggingFacePipeline from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline tokenizer = AutoTokenizer.from_pretrained("google/flan-t5-xl") model = AutoModelForSeq2SeqLM.from_pretrained("google/flan-t5-xl") pipe = pipeline( "text2text-generation", model=model, tokenizer=tokenizer, max_length=64, temperature=0 ) llm = HuggingFacePipeline(pipeline=pipe) llm_chain = LLMChain(prompt=prompt, llm=llm) print(llm_chain.run(question)) - 方案4:检查网络稳定性
确保网络能稳定访问HuggingFace推理服务,必要时切换网络环境。
内容的提问来源于stack exchange,提问作者Jennie94
相关产品推荐
相关产品推荐

