使用LangChain+HuggingFaceEndpoint调用flan-t5-large时遇StopIteration错误
问题描述
我尝试在LangChain流水线中使用langchain_huggingface.HuggingFaceEndpoint集成调用Hugging Face的"google/flan-t5-large"模型,代码如下:
from langchain.prompts import PromptTemplate from langchain_huggingface import HuggingFaceEndpoint from langchain_core.runnables import RunnableSequence import os # Set your API token os.environ['HUGGINGFACEHUB_API_TOKEN'] = 'hf_***************' # Initialize the updated Hugging Face model flan_t5 = HuggingFaceEndpoint( repo_id="google/flan-t5-large", temperature=1e-10, task="text2text-generation" ) # Define the prompt template template = "Translate the following to French:\n\n{question}" prompt = PromptTemplate(template=template, input_variables=["question"]) # Create the RunnableSequence chain chain = prompt | flan_t5 # Run the chain with input question = "Hello" response = chain.invoke({"question": question}) print("Output:", response)
但出现如下错误:
/usr/local/lib/python3.11/dist-packages/huggingface_hub/inference/_providers/__init__.py in get_provider_helper(provider, task, model) 189 raise ValueError("Specifying a model is required when provider is 'auto'") 190 provider_mapping = _fetch_inference_provider_mapping(model) --> 191 provider = next(iter(provider_mapping)).provider StopIteration
已尝试的操作:
- 验证Hugging Face令牌有效
- 确认"google/flan-t5-large"存在且支持text2text-generation
- 尝试其他仓库ID(如google/flan-t5-xl),仍出现相同问题
疑问:
- 为何会出现StopIteration错误?如何修复?
- 初始化HuggingFaceEndpoint或提供者推断方式是否存在问题?
问题分析与解决方案
错误原因
StopIteration错误的根源是Hugging Face Hub的推理提供者自动推断失败:当未指定provider参数时,HuggingFaceEndpoint默认用auto模式获取模型的可用推理提供者,但google/flan-t5-large这类模型没有配置官方托管的推理端点,或者当前无法获取到有效提供者列表,导致provider_mapping为空,调用next(iter(provider_mapping))时触发迭代器耗尽的错误。
修复方案
方案1:显式指定自定义部署的推理端点
如果你已经在Hugging Face上部署了自己的推理端点,初始化HuggingFaceEndpoint时直接指定endpoint_url参数,跳过自动提供者推断:
flan_t5 = HuggingFaceEndpoint( endpoint_url="https://<your-endpoint-id>.hf.space", # 替换为自己的部署端点URL temperature=1e-10, task="text2text-generation" )
方案2:本地运行模型(无需依赖远程端点)
改用langchain_huggingface.HuggingFacePipeline直接本地加载模型运行,避免远程端点依赖:
from langchain_huggingface import HuggingFacePipeline from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline # 加载本地模型和tokenizer model_id = "google/flan-t5-large" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForSeq2SeqLM.from_pretrained(model_id) # 创建text2text-generation pipeline pipe = pipeline( "text2text-generation", model=model, tokenizer=tokenizer, temperature=1e-10, max_new_tokens=200 ) # 初始化LangChain的HuggingFacePipeline flan_t5 = HuggingFacePipeline(pipeline=pipe) # 后续prompt和chain代码保持不变 template = "Translate the following to French:\n\n{question}" prompt = PromptTemplate(template=template, input_variables=["question"]) chain = prompt | flan_t5 question = "Hello" response = chain.invoke({"question": question}) print("Output:", response)
方案3:显式指定Hugging Face Inference API提供者
如果模型支持Hugging Face免费Inference API,可显式指定provider="huggingface",确保API令牌有访问权限:
flan_t5 = HuggingFaceEndpoint( repo_id="google/flan-t5-large", temperature=1e-10, task="text2text-generation", provider="huggingface" )
注意:若模型未开放免费Inference API访问,此方法会失败,建议优先选择本地运行方案。
初始化HuggingFaceEndpoint的注意事项
- 默认的
auto提供者推断仅适用于有官方托管推理端点的模型,大部分开源模型需要自行部署端点或本地运行。 - 依赖远程推理时,必须确保模型已部署到Hugging Face Inference Endpoints,并正确配置
endpoint_url或指定有效provider。
内容的提问来源于stack exchange,提问作者coderbhai
相关产品推荐
相关产品推荐

