You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Hugging Face Transformers中的Tokenizer导出至CoreML?

将Hugging Face Tokenizer适配CoreML的实现方案

Tokenizer本质是文本预处理工具,并非神经网络模型,因此没有直接“导出为CoreML模型”的方法,需将其预处理逻辑与CoreML模型整合,以下是两种实用方案:

方案一:构建包含Tokenizer逻辑的CoreML Pipeline(适合Python环境推理)

该方案通过CoreML Tools将Tokenizer的预处理逻辑封装为自定义层,与已导出的BERT CoreML模型整合为Pipeline,实现输入文本直接得到推理结果。

步骤1:安装依赖

pip install coremltools transformers torch

步骤2:编写整合代码

import coremltools as ct
from transformers import AutoTokenizer

# 加载已导出的BERT CoreML模型(替换为你的模型路径)
bert_coreml_model = ct.models.MLModel("bert_token_classification_coreml.mlmodel")

# 加载Tokenizer
tokenizer = AutoTokenizer.from_pretrained("huawei-noah/TinyBERT_General_4L_312D")
max_seq_len = tokenizer.model_max_length

# 定义文本预处理类,封装Tokenizer逻辑
class TextTokenizerLayer(ct.CustomPythonLayer):
    def __init__(self, tokenizer):
        self.tokenizer = tokenizer

    def compute(self, input_dict):
        # 获取输入文本
        input_text = input_dict["input_text"].item()
        # 执行分词、截断、填充
        tokenized = self.tokenizer(
            input_text,
            return_tensors="np",
            truncation=True,
            padding="max_length",
            max_length=max_seq_len
        )
        # 返回模型所需的输入张量
        return {
            "input_ids": tokenized["input_ids"].squeeze(0),
            "attention_mask": tokenized["attention_mask"].squeeze(0)
        }

# 创建预处理模型
input_spec = ct.TensorType(name="input_text", shape=(1,), dtype=str)
preprocessor = ct.models.Model(
    ct.CustomPythonLayer(
        name="text_tokenizer",
        input_features=[input_spec],
        output_features=[
            ct.TensorType(name="input_ids", shape=(max_seq_len,), dtype=int),
            ct.TensorType(name="attention_mask", shape=(max_seq_len,), dtype=int)
        ],
        python_module=TextTokenizerLayer(tokenizer)
    )
)

# 构建完整Pipeline
pipeline = ct.models.pipeline.Pipeline(
    input_features=[input_spec],
    output_features=bert_coreml_model.output_features,
    steps=[("tokenizer", preprocessor), ("bert_model", bert_coreml_model)]
)

# 保存Pipeline模型
pipeline.save("tinybert_tokenizer_coreml_pipeline.mlmodel")

使用说明

保存后的Pipeline模型可直接接收文本输入,自动完成分词和模型推理:

import coremltools as ct

pipeline_model = ct.models.MLModel("tinybert_tokenizer_coreml_pipeline.mlmodel")
result = pipeline_model.predict({"input_text": "Hugging Face is creating a tool that democratizes AI."})
print(result)

方案二:导出Tokenizer配置,在移动端实现原生分词逻辑(适合iOS/macOS部署)

若需在iOS/macOS端部署,由于CoreML不支持Python自定义层,需导出Tokenizer的核心配置,用Swift等原生代码实现分词逻辑,再传入CoreML模型。

步骤1:导出Tokenizer配置

from transformers import AutoTokenizer
import json

tokenizer = AutoTokenizer.from_pretrained("huawei-noah/TinyBERT_General_4L_312D")

# 导出词汇表
tokenizer.save_vocabulary("./tinybert_tokenizer")
# 导出Tokenizer配置(含特殊标记、最大序列长度等)
with open("./tinybert_tokenizer/config.json", "w") as f:
    json.dump(tokenizer.get_config(), f)

步骤2:在Swift中实现分词逻辑

将导出的词汇表和配置文件添加到iOS项目中,编写Swift代码实现类似Tokenizer的功能,包括:

  • 读取词汇表和特殊标记(如[CLS]、[SEP])
  • 对输入文本进行分词、子词拆分
  • 截断或填充到指定的最大序列长度
  • 生成input_ids和attention_mask张量,传入CoreML模型

示例Swift代码片段

import CoreML

// 加载Tokenizer配置和词汇表
let configPath = Bundle.main.path(forResource: "config", ofType: "json")!
let vocabPath = Bundle.main.path(forResource: "vocab", ofType: "txt")!
// 实现分词逻辑(省略具体实现,需根据Tokenizer类型编写)
func tokenize(text: String) -> (inputIds: [Int32], attentionMask: [Int32]) {
    // 此处编写分词、截断、填充逻辑
    return (inputIds: [], attentionMask: [])
}

// 加载CoreML模型
let bertModel = TinyBERT_CoreML()
// 处理文本
let (inputIds, attentionMask) = tokenize(text: "Hugging Face is creating a tool that democratizes AI.")
// 转换为模型所需的MLMultiArray
let inputIdsArray = try! MLMultiArray(shape: [1, 312], dataType: .int32)
let attentionMaskArray = try! MLMultiArray(shape: [1, 312], dataType: .int32)
// 填充数据(省略)
// 执行推理
let result = try! bertModel.prediction(input_ids: inputIdsArray, attention_mask: attentionMaskArray)

关键说明

  • Tokenizer的核心是文本预处理逻辑,而非可直接转换的模型,需根据部署场景选择合适的整合方式。
  • 若使用Pipeline方案,仅适合Python环境下的推理;移动端部署需采用原生代码实现分词逻辑。

内容的提问来源于stack exchange,提问作者Franck Dernoncourt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 22:04:57