You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CPU本地加载GPT-OSS-20B模型时遭遇KeyError问题求助

CPU本地加载GPT-OSS-20B模型时遭遇KeyError问题求助

我现在尝试在纯CPU环境下用Hugging Face Transformers本地加载gpt-oss-20B模型,但遇到了一个KeyError问题,折腾了好久没解决,来求助大家!

我的最小复现代码

from transformers import pipeline
model_path = "/mnt/d/Projects/models/gpt-oss-20b"
pipe = pipeline("text-generation", model=model_path, torch_dtype="auto", device_map="auto")
pipe("Hello", max_new_tokens=20)

报错信息及关键日志

运行后首先输出了这些提示:

Using MXFP4 quantized models requires a GPU, we will default to dequantizing the model to bf16
Loading checkpoint shards: 100%
Some parameters are on the meta device because they were offloaded to the cpu and disk.
Device set to use cpu

然后就抛出了KeyError:

Traceback (most recent call last):
  File "/home/dev/projects/wolf-in-ai-clothing/convo_test.py", line 19, in invoke
    response = model(user_message, max_new_tokens=20, num_return_sequences=1)
  File ".../transformers/pipelines/text_generation.py", line 419, in _forward
    output = self.model.generate(input_ids=input_ids, attention_mask=attention_mask, **generate_kwargs)
  File ".../transformers/models/gpt_oss/modeling_gpt_oss.py", line 375, in forward
    hidden_states, _ = self.mlp(hidden_states) # diff with llama: router scores
  File ".../transformers/models/gpt_oss/modeling_gpt_oss.py", line 159, in forward
    routed_out = self.experts(hidden_states, router_indices=router_indices, routing_weights=router_scores)
  File ".../accelerate/utils/offload.py", line 118, in __getitem__
    return self.dataset[f"{self.prefix}{key}"]
  File ".../accelerate/utils/offload.py", line 165, in __getitem__
    weight_info = self.index[key]
KeyError: 'model.layers.5.mlp.experts.gate_up_proj'

我已经确认模型目录是存在的,里面的模型文件也都齐全。之前在Hugging Face论坛看到过类似问题,跟着@noobaymax的建议试了这些操作:

  • 安装triton主分支版本:pip install git+https://github.com/triton-lang/triton.git@main#subdirectory=python/triton_kernels
  • 安装transformers的最新开发版:pip install git+https://github.com/huggingface/transformers.git
  • 安装相关kernel依赖(跟着步骤执行,但未明确具体包名)

但执行完这些后,问题还是完全一样,没有任何改善。

我的环境信息

  • Python:3.12.3
  • Transformers:4.56.0.dev0(也试过稳定版4.55.1)
  • PyTorch:2.8.0
  • Accelerate:1.10.0
  • 系统:WSL2上的Ubuntu 22.04,无GPU,32GB RAM

有没有大佬遇到过类似的问题?或者知道纯CPU环境下正确加载这个模型的方法?麻烦指点一下!


内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 07:53:06