在Google Colab运行Hugging Face BERT词性标注模型报错求助
问题解决方法
核心问题分析
- 模型类型不兼容:你用
AutoModelWithHeads加载的模型属于BertModelWithHeads类型,原生transformers的token-classificationpipeline不支持该类型,仅支持列表中列出的BertAdapterModel这类适配过的模型。 - 标签映射缺失:加载适配器后,模型主配置的
id2label未同步适配器的标签映射,导致预测时找不到对应ID的标签,抛出KeyError:16。 - 环境差异:Hugging Face API已配置好适配适配器的运行环境,而你的Colab环境用的是原生
transformers,未适配AdapterHub的模型结构。
修正后的代码方案
步骤1:安装适配AdapterHub的库
先在Colab中安装官方支持AdapterHub的分支库:
!pip install adapter-transformers
步骤2:使用正确的模型类加载并运行
替换原有代码为以下内容:
from adapter_transformers import AutoAdapterModel, AutoTokenizer from adapter_transformers import pipeline # 用AutoAdapterModel加载基础模型 model = AutoAdapterModel.from_pretrained("bert-base-uncased") # 加载词性标注适配器,同时自动激活并同步标签配置 adapter_name = model.load_adapter("AdapterHub/bert-base-uncased-pf-ud_pos", source="hf", set_active=True) # 加载对应的tokenizer tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased") # 初始化token分类pipeline,此时模型类型为BertAdapterModel,属于支持列表内 token_classification = pipeline( "token-classification", model=model, tokenizer=tokenizer, aggregation_strategy="NONE" ) # 测试预测 res = token_classification("Take out the trash bag from the bin and replace it.") print(res)
备选方案:直接加载完整预训练模型
如果不想分开加载基础模型和适配器,也可以直接加载AdapterHub上的完整模型(已包含基础模型和适配器):
from adapter_transformers import AutoModelForTokenClassification, AutoTokenizer from adapter_transformers import pipeline model = AutoModelForTokenClassification.from_pretrained("AdapterHub/bert-base-uncased-pf-ud_pos") tokenizer = AutoTokenizer.from_pretrained("AdapterHub/bert-base-uncased-pf-ud_pos") token_classification = pipeline( "token-classification", model=model, tokenizer=tokenizer, aggregation_strategy="NONE" ) res = token_classification("Take out the trash bag from the bin and replace it.") print(res)
为什么Hugging Face API能正常运行?
Hugging Face API的后台环境已预配置adapter-transformers库,并且对AdapterHub的模型做了适配处理,自动兼容模型类型和标签映射,所以无需额外配置即可运行。
内容的提问来源于stack exchange,提问作者Encipher
相关产品推荐
相关产品推荐

