You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LangChain+HuggingFaceEndpoint调用flan-t5-large时遇StopIteration错误

问题描述

我尝试在LangChain流水线中使用langchain_huggingface.HuggingFaceEndpoint集成调用Hugging Face的"google/flan-t5-large"模型,代码如下:

from langchain.prompts import PromptTemplate
from langchain_huggingface import HuggingFaceEndpoint
from langchain_core.runnables import RunnableSequence
import os

# Set your API token
os.environ['HUGGINGFACEHUB_API_TOKEN'] = 'hf_***************'

# Initialize the updated Hugging Face model
flan_t5 = HuggingFaceEndpoint(
    repo_id="google/flan-t5-large",
    temperature=1e-10,
    task="text2text-generation"
)

# Define the prompt template
template = "Translate the following to French:\n\n{question}"
prompt = PromptTemplate(template=template, input_variables=["question"])

# Create the RunnableSequence chain
chain = prompt | flan_t5

# Run the chain with input
question = "Hello"
response = chain.invoke({"question": question})

print("Output:", response)

但出现如下错误:

/usr/local/lib/python3.11/dist-packages/huggingface_hub/inference/_providers/__init__.py in get_provider_helper(provider, task, model)
    189             raise ValueError("Specifying a model is required when provider is 'auto'")
    190         provider_mapping = _fetch_inference_provider_mapping(model)
--> 191         provider = next(iter(provider_mapping)).provider
StopIteration

已尝试的操作:

  • 验证Hugging Face令牌有效
  • 确认"google/flan-t5-large"存在且支持text2text-generation
  • 尝试其他仓库ID(如google/flan-t5-xl),仍出现相同问题

疑问:

  • 为何会出现StopIteration错误?如何修复?
  • 初始化HuggingFaceEndpoint或提供者推断方式是否存在问题?
问题分析与解决方案

错误原因

StopIteration错误的根源是Hugging Face Hub的推理提供者自动推断失败:当未指定provider参数时,HuggingFaceEndpoint默认用auto模式获取模型的可用推理提供者,但google/flan-t5-large这类模型没有配置官方托管的推理端点,或者当前无法获取到有效提供者列表,导致provider_mapping为空,调用next(iter(provider_mapping))时触发迭代器耗尽的错误。

修复方案

方案1:显式指定自定义部署的推理端点

如果你已经在Hugging Face上部署了自己的推理端点,初始化HuggingFaceEndpoint时直接指定endpoint_url参数,跳过自动提供者推断:

flan_t5 = HuggingFaceEndpoint(
    endpoint_url="https://<your-endpoint-id>.hf.space",  # 替换为自己的部署端点URL
    temperature=1e-10,
    task="text2text-generation"
)

方案2:本地运行模型(无需依赖远程端点)

改用langchain_huggingface.HuggingFacePipeline直接本地加载模型运行,避免远程端点依赖:

from langchain_huggingface import HuggingFacePipeline
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM, pipeline

# 加载本地模型和tokenizer
model_id = "google/flan-t5-large"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id)

# 创建text2text-generation pipeline
pipe = pipeline(
    "text2text-generation",
    model=model,
    tokenizer=tokenizer,
    temperature=1e-10,
    max_new_tokens=200
)

# 初始化LangChain的HuggingFacePipeline
flan_t5 = HuggingFacePipeline(pipeline=pipe)

# 后续prompt和chain代码保持不变
template = "Translate the following to French:\n\n{question}"
prompt = PromptTemplate(template=template, input_variables=["question"])
chain = prompt | flan_t5

question = "Hello"
response = chain.invoke({"question": question})
print("Output:", response)

方案3:显式指定Hugging Face Inference API提供者

如果模型支持Hugging Face免费Inference API,可显式指定provider="huggingface",确保API令牌有访问权限:

flan_t5 = HuggingFaceEndpoint(
    repo_id="google/flan-t5-large",
    temperature=1e-10,
    task="text2text-generation",
    provider="huggingface"
)

注意:若模型未开放免费Inference API访问,此方法会失败,建议优先选择本地运行方案。

初始化HuggingFaceEndpoint的注意事项

  • 默认的auto提供者推断仅适用于有官方托管推理端点的模型,大部分开源模型需要自行部署端点或本地运行。
  • 依赖远程推理时,必须确保模型已部署到Hugging Face Inference Endpoints,并正确配置endpoint_url或指定有效provider。

内容的提问来源于stack exchange,提问作者coderbhai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 15:17:33