如何让导出至CoreML的BERT类PyTorch模型支持任意文本输入?
问题
使用下方代码将基于BERT的PyTorch模型导出至CoreML时,因导出阶段使用了固定的dummy输入:
dummy_input = tokenizer("A French fan", return_tensors="pt")
导致在macOS上测试时,CoreML模型仅能处理该特定输入。测试其他文本(如predict("A football fan is standing in the stadium."))时触发错误:
NSLocalizedDescription = "MultiArray shape (1 x 12) does not match the shape (1 x 5) specified in the model description";
导出脚本
# -*- coding: utf-8 -*- """Core ML Export pip install transformers torch coremltools nltk """ import os from transformers import AutoModelForTokenClassification, AutoTokenizer import torch import torch.nn as nn import nltk import coremltools as ct nltk.download('punkt') # Load the model and tokenizer model_path = os.path.join('model') model = AutoModelForTokenClassification.from_pretrained(model_path, local_files_only=True) tokenizer = AutoTokenizer.from_pretrained(model_path, local_files_only=True) # Modify the model's forward method to return a tuple class ModifiedModel(nn.Module): def __init__(self, model): super(ModifiedModel, self).__init__() self.model = model self.device = model.device # Add the device attribute def forward(self, input_ids, attention_mask, token_type_ids=None): outputs = self.model(input_ids=input_ids, attention_mask=attention_mask, token_type_ids=token_type_ids) return outputs.logits modified_model = ModifiedModel(model) # Export to Core ML def convert_to_coreml(model, tokenizer): # Define a dummy input for tracing dummy_input = tokenizer("A French fan", return_tensors="pt") dummy_input = {k: v.to(model.device) for k, v in dummy_input.items()} # Trace the model with the dummy input traced_model = torch.jit.trace(model, ( dummy_input['input_ids'], dummy_input['attention_mask'], dummy_input.get('token_type_ids'))) # Convert to Core ML inputs = [ ct.TensorType(name="input_ids", shape=dummy_input['input_ids'].shape), ct.TensorType(name="attention_mask", shape=dummy_input['attention_mask'].shape) ] if 'token_type_ids' in dummy_input: inputs.append(ct.TensorType(name="token_type_ids", shape=dummy_input['token_type_ids'].shape)) mlmodel = ct.convert(traced_model, inputs=inputs) # Save the Core ML model mlmodel.save("model.mlmodel") print("Model exported to Core ML successfully") convert_to_coreml(modified_model, tokenizer)
预测脚本
import os from transformers import AutoModelForTokenClassification, AutoTokenizer import torch import torch.nn as nn import nltk import coremltools as ct from coremltools.models import MLModel import numpy as np from transformers import AutoTokenizer import nltk nltk.download('punkt') # Load the Core ML model model = MLModel('model.mlmodel') # Load the tokenizer model_path = 'model' tokenizer = AutoTokenizer.from_pretrained(model_path, local_files_only=True) def prepare_input(text, tokenizer): tokens = nltk.tokenize.word_tokenize(text) tokenized_inputs = tokenizer(tokens, is_split_into_words=True, return_tensors="np") input_ids = tokenized_inputs['input_ids'].astype(np.int32) attention_mask = tokenized_inputs['attention_mask'].astype(np.int32) input_data = { 'input_ids': input_ids, 'attention_mask': attention_mask } if 'token_type_ids' in tokenized_inputs: input_data['token_type_ids'] = tokenized_inputs['token_type_ids'].astype(np.int32) return input_data, tokens def predict(text): # Prepare the input input_data, tokens = prepare_input(text, tokenizer) # Make the prediction prediction = model.predict(input_data) # Extract the predicted labels logits = prediction['output'] # Adjust this key according to your model's output predicted_label = np.argmax(logits, axis=-1)[0] # Display the results for word, label in zip(tokens, predicted_label): print(f"{word}: {model.model_description.outputDescriptions[0].dictionaryType.int64KeyType.stringDictionary[label]}") # Test the model with a sentence predict("A French fan")
环境信息
- 导出脚本:Ubuntu 20.04系统、Python 3.10、torch 2.3.1(Windows 10环境无法运行)
- 预测脚本:需在macOS 10.13及以上系统运行
解决方案
问题根源是导出时用了固定形状的dummy input,导致CoreML模型被固化为只能处理该长度的输入。要支持任意文本输入,需修改导出脚本,将输入形状设置为动态维度:
核心修改步骤
修改convert_to_coreml函数,关键是将输入序列维度设为None(表示可变长度),同时使用CoreML的MLProgram格式提升动态形状支持:
def convert_to_coreml(model, tokenizer): # 创建动态dummy输入:batch_size固定为1,序列长度设为任意值仅用于模型追踪 batch_size = 1 seq_len = 10 dummy_input_ids = torch.randint(low=0, high=tokenizer.vocab_size, size=(batch_size, seq_len)).to(model.device) dummy_attention_mask = torch.ones((batch_size, seq_len)).to(model.device) dummy_token_type_ids = torch.zeros((batch_size, seq_len)).to(model.device) if tokenizer.bos_token_id else None # 整理dummy输入字典 dummy_input = { 'input_ids': dummy_input_ids, 'attention_mask': dummy_attention_mask } if dummy_token_type_ids is not None: dummy_input['token_type_ids'] = dummy_token_type_ids # 追踪模型,关闭强制校验以适配动态形状 traced_model = torch.jit.trace(model, ( dummy_input['input_ids'], dummy_input['attention_mask'], dummy_input.get('token_type_ids')), check_trace=False) # 定义CoreML输入:将序列维度设为None,标记为可变长度 inputs = [ ct.TensorType(name="input_ids", shape=(batch_size, None)), ct.TensorType(name="attention_mask", shape=(batch_size, None)) ] if 'token_type_ids' in dummy_input: inputs.append(ct.TensorType(name="token_type_ids", shape=(batch_size, None))) # 转换为CoreML模型,指定使用MLProgram格式 mlmodel = ct.convert( traced_model, inputs=inputs, convert_to="mlprogram" ) # 保存模型 mlmodel.save("model.mlmodel") print("Model exported to Core ML successfully with dynamic input support")
关键说明
- 动态维度设置:将输入形状的序列维度设为
None,告知CoreML该维度可接收任意长度(不超过BERT模型的最大序列长度,如512)。 - MLProgram格式:该格式比旧的NeuralNetwork格式更适配现代PyTorch特性,对动态形状的支持更稳定。
- dummy输入作用:仅用于模型追踪,不限制实际输入长度。
验证
重新导出模型后,使用原预测脚本测试任意长度文本,即可避免形状不匹配的错误。
内容的提问来源于stack exchange,提问作者Franck Dernoncourt
相关产品推荐
相关产品推荐

