解决Colab中ModuleNotFoundError: No module named 'transformers.models.mmbt'报错
解决ModuleNotFoundError: No module named 'transformers.models.mmbt'
问题背景
在Google Colab运行基于classla/xlm-roberta-base-multilingual-text-genre-classifier的文本分类代码时,突然出现上述报错,此前代码可正常运行。当前环境:
- transformers 4.30.2
- simpletransformers 0.63.11
- Python 3
原因
simpletransformers 0.63.11版本的多模态分类模块依赖transformers.models.mmbt,但transformers从4.29.0版本开始,已将MMBT(多模态双向Transformer)从核心库中移除,移至单独的扩展组件。
解决方案
方案1:安装transformers的MMBT扩展组件
在Colab中执行以下命令,重新安装带MMBT支持的指定版本transformers:
!pip uninstall -y transformers !pip install transformers[mmbt]==4.30.2
方案2:降级transformers到仍包含MMBT的版本
如果方案1无效,可降级transformers到4.28.1版本(该版本仍将MMBT保留在核心库中):
!pip uninstall -y transformers !pip install transformers==4.28.1
方案3:绕过simpletransformers,直接使用transformers核心库
因为实际只用到XLM-Roberta的文本分类功能,完全可以不用simpletransformers,直接用transformers原生接口实现,避免依赖问题。替换代码如下:
from transformers import AutoTokenizer, AutoModelForSequenceClassification import torch # 加载模型和分词器 tokenizer = AutoTokenizer.from_pretrained("classla/xlm-roberta-base-multilingual-text-genre-classifier") model = AutoModelForSequenceClassification.from_pretrained("classla/xlm-roberta-base-multilingual-text-genre-classifier") model.to("cuda") # 使用GPU # 待预测文本 texts = [ "How to create a good text classification model? First step is to prepare good data. Make sure not to skip the exploratory data analysis. Pre-process the text if necessary for the task. The next step is to perform hyperparameter search to find the optimum hyperparameters. After fine-tuning the model, you should look into the predictions and analyze the model's performance. You might want to perform the post-processing of data as well and keep only reliable predictions.", "On our site, you can find a great genre identification model which you can use for thousands of different tasks. With our model, you can fastly and reliably obtain high-quality genre predictions and explore which genres exist in your corpora. Available for free!" ] # 预处理文本 inputs = tokenizer(texts, padding=True, truncation=True, max_length=512, return_tensors="pt").to("cuda") # 预测 with torch.no_grad(): outputs = model(**inputs) predictions = torch.argmax(outputs.logits, dim=-1).cpu().numpy() # 获取标签 labels = [model.config.id2label[i] for i in predictions] print(predictions) # 输出: [3 8] print(labels) # 输出: ['Instruction', 'Promotion']
内容的提问来源于stack exchange,提问作者simKO
相关产品推荐
相关产品推荐

