如何将Hugging Face Transformers中的Tokenizer导出至CoreML?
将Hugging Face Tokenizer适配CoreML的实现方案
Tokenizer本质是文本预处理工具,并非神经网络模型,因此没有直接“导出为CoreML模型”的方法,需将其预处理逻辑与CoreML模型整合,以下是两种实用方案:
方案一:构建包含Tokenizer逻辑的CoreML Pipeline(适合Python环境推理)
该方案通过CoreML Tools将Tokenizer的预处理逻辑封装为自定义层,与已导出的BERT CoreML模型整合为Pipeline,实现输入文本直接得到推理结果。
步骤1:安装依赖
pip install coremltools transformers torch
步骤2:编写整合代码
import coremltools as ct from transformers import AutoTokenizer # 加载已导出的BERT CoreML模型(替换为你的模型路径) bert_coreml_model = ct.models.MLModel("bert_token_classification_coreml.mlmodel") # 加载Tokenizer tokenizer = AutoTokenizer.from_pretrained("huawei-noah/TinyBERT_General_4L_312D") max_seq_len = tokenizer.model_max_length # 定义文本预处理类,封装Tokenizer逻辑 class TextTokenizerLayer(ct.CustomPythonLayer): def __init__(self, tokenizer): self.tokenizer = tokenizer def compute(self, input_dict): # 获取输入文本 input_text = input_dict["input_text"].item() # 执行分词、截断、填充 tokenized = self.tokenizer( input_text, return_tensors="np", truncation=True, padding="max_length", max_length=max_seq_len ) # 返回模型所需的输入张量 return { "input_ids": tokenized["input_ids"].squeeze(0), "attention_mask": tokenized["attention_mask"].squeeze(0) } # 创建预处理模型 input_spec = ct.TensorType(name="input_text", shape=(1,), dtype=str) preprocessor = ct.models.Model( ct.CustomPythonLayer( name="text_tokenizer", input_features=[input_spec], output_features=[ ct.TensorType(name="input_ids", shape=(max_seq_len,), dtype=int), ct.TensorType(name="attention_mask", shape=(max_seq_len,), dtype=int) ], python_module=TextTokenizerLayer(tokenizer) ) ) # 构建完整Pipeline pipeline = ct.models.pipeline.Pipeline( input_features=[input_spec], output_features=bert_coreml_model.output_features, steps=[("tokenizer", preprocessor), ("bert_model", bert_coreml_model)] ) # 保存Pipeline模型 pipeline.save("tinybert_tokenizer_coreml_pipeline.mlmodel")
使用说明
保存后的Pipeline模型可直接接收文本输入,自动完成分词和模型推理:
import coremltools as ct pipeline_model = ct.models.MLModel("tinybert_tokenizer_coreml_pipeline.mlmodel") result = pipeline_model.predict({"input_text": "Hugging Face is creating a tool that democratizes AI."}) print(result)
方案二:导出Tokenizer配置,在移动端实现原生分词逻辑(适合iOS/macOS部署)
若需在iOS/macOS端部署,由于CoreML不支持Python自定义层,需导出Tokenizer的核心配置,用Swift等原生代码实现分词逻辑,再传入CoreML模型。
步骤1:导出Tokenizer配置
from transformers import AutoTokenizer import json tokenizer = AutoTokenizer.from_pretrained("huawei-noah/TinyBERT_General_4L_312D") # 导出词汇表 tokenizer.save_vocabulary("./tinybert_tokenizer") # 导出Tokenizer配置(含特殊标记、最大序列长度等) with open("./tinybert_tokenizer/config.json", "w") as f: json.dump(tokenizer.get_config(), f)
步骤2:在Swift中实现分词逻辑
将导出的词汇表和配置文件添加到iOS项目中,编写Swift代码实现类似Tokenizer的功能,包括:
- 读取词汇表和特殊标记(如
[CLS]、[SEP]) - 对输入文本进行分词、子词拆分
- 截断或填充到指定的最大序列长度
- 生成
input_ids和attention_mask张量,传入CoreML模型
示例Swift代码片段
import CoreML // 加载Tokenizer配置和词汇表 let configPath = Bundle.main.path(forResource: "config", ofType: "json")! let vocabPath = Bundle.main.path(forResource: "vocab", ofType: "txt")! // 实现分词逻辑(省略具体实现,需根据Tokenizer类型编写) func tokenize(text: String) -> (inputIds: [Int32], attentionMask: [Int32]) { // 此处编写分词、截断、填充逻辑 return (inputIds: [], attentionMask: []) } // 加载CoreML模型 let bertModel = TinyBERT_CoreML() // 处理文本 let (inputIds, attentionMask) = tokenize(text: "Hugging Face is creating a tool that democratizes AI.") // 转换为模型所需的MLMultiArray let inputIdsArray = try! MLMultiArray(shape: [1, 312], dataType: .int32) let attentionMaskArray = try! MLMultiArray(shape: [1, 312], dataType: .int32) // 填充数据(省略) // 执行推理 let result = try! bertModel.prediction(input_ids: inputIdsArray, attention_mask: attentionMaskArray)
关键说明
- Tokenizer的核心是文本预处理逻辑,而非可直接转换的模型,需根据部署场景选择合适的整合方式。
- 若使用Pipeline方案,仅适合Python环境下的推理;移动端部署需采用原生代码实现分词逻辑。
内容的提问来源于stack exchange,提问作者Franck Dernoncourt
相关产品推荐
相关产品推荐

