You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Transformers库LlamaForCausalLM的base_model无限递归问题咨询

Transformers Pipeline中LlamaForCausalLM的base_model无限递归问题

使用transformers库的pipeline初始化时,访问LlamaForCausalLM模型的base_model属性出现无限递归现象。

示例代码

from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline

# let's rock the smallest model on hf
MODEL_NAME = "arnir0/Tiny-LLM"

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)

INPUT = ["the secret to baking a really good cake is ", "a baguette is "]

if __name__ == "__main__":
    pipeline = pipeline(
        task="text-generation",
        model=MODEL_NAME,
        device="mps"
    )
    for out in pipeline(INPUT):
        print(out[0]["generated_text"])

调试现象

在VSCode调试时发现:创建text-generation类型的pipeline后,其model属性为LlamaForCausalLM实例,该实例的base_model指向LlamaModel;但访问LlamaModel的base_model属性时,会无限递归,始终指向自身对象,交互示例如下:

>>> from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
>>> MODEL_NAME = "arnir0/Tiny-LLM"
>>> tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
>>> model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)
>>> INPUT = ["the secret to baking a really good cake is ", "a baguette is "]
>>> pipeline = pipeline(task="text-generation", model=MODEL_NAME, device="mps")
Device set to use mps
>>> model = pipeline.model
>>> model.base_model
LlamaModel(
  (embed_tokens): Embedding(32000, 192)
  (layers): ModuleList(
    (0): LlamaDecoderLayer(
      (self_attn): LlamaAttention(
        (q_proj): Linear(in_features=192, out_features=192, bias=False)
        (k_proj): Linear(in_features=192, out_features=96, bias=False)
        (v_proj): Linear(in_features=192, out_features=96, bias=False)
        (o_proj): Linear(in_features=192, out_features=192, bias=False)
      )
      (mlp): LlamaMLP(
        (gate_proj): Linear(in_features=192, out_features=1024, bias=False)
        (up_proj): Linear(in_features=192, out_features=1024, bias=False)
        (down_proj): Linear(in_features=1024, out_features=192, bias=False)
        (act_fn): SiLUActivation()
      )
      (input_layernorm): LlamaRMSNorm((192,), eps=1e-05)
      (post_attention_layernorm): LlamaRMSNorm((192,), eps=1e-05)
    )
  )
  (norm): LlamaRMSNorm((192,), eps=1e-05)
  (rotary_emb): LlamaRotaryEmbedding()
)
>>> model.base_model.base_model
LlamaModel(
  (embed_tokens): Embedding(32000, 192)
  (layers): ModuleList(
    (0): LlamaDecoderLayer(
      (self_attn): LlamaAttention(
        (q_proj): Linear(in_features=192, out_features=192, bias=False)
        (k_proj): Linear(in_features=192, out_features=96, bias=False)
        (v_proj): Linear(in_features=192, out_features=96, bias=False)
        (o_proj): Linear(in_features=192, out_features=192, bias=False)
      )
      (mlp): LlamaMLP(
        (gate_proj): Linear(in_features=192, out_features=1024, bias=False)
        (up_proj): Linear(in_features=192, out_features=1024, bias=False)
        (down_proj): Linear(in_features=1024, out_features=192, bias=False)
        (act_fn): SiLUActivation()
      )
      (input_layernorm): LlamaRMSNorm((192,), eps=1e-05)
      (post_attention_layernorm): LlamaRMSNorm((192,), eps=1e-05)
    )
  )
  (norm): LlamaRMSNorm((192,), eps=1e-05)
  (rotary_emb): LlamaRotaryEmbedding()
)

问题

  1. 该递归循环是否会终止?是否存在终止条件?
  2. 为何model属性会陷入此循环?是transformers库的设计选择还是有特定原因?
  3. 该递归行为是否与model.config相关?我检查过config看似正常,但不确定二者关联。

环境信息

  • Python 3.13
  • transformers 4.57.1
  • torch 2.9.1
  • sentencepiece 0.2.1
  • tokenizers 0.22.1
  • uv 0.7.22
  • VSCode: 1.106.0, arm64

解答

  1. 这个递归不会自动终止,没有内置终止条件。因为LlamaModel的base_model属性被设计为返回自身,只要持续链式调用.base_model,就会一直指向同一个LlamaModel实例,直到手动停止访问。

  2. 这是transformers库的刻意设计选择:

    • LlamaForCausalLM是带有语言模型头部(LM Head)的封装类,它的base_model指向核心的LlamaModel(不带预测头部的模型主体结构)。
    • LlamaModel作为最底层的基础模型结构,它的base_model属性返回自身,以此统一接口——不管用户拿到的是带头部的封装模型还是基础模型,调用.base_model都能获取到最核心的模型主体,无需区分不同层级的模型类结构。
  3. 该递归行为和model.config无关。config主要存储模型的架构参数(如维度、层数、注意力头数等),而base_model是模型类的属性设计逻辑,属于类结构层面的实现,和配置信息没有直接关联。你的config正常是合理的,这个递归现象完全是库的类结构设计导致的,不是配置异常。

内容的提问来源于stack exchange,提问作者Hank Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 01:29:51