You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速M1 Pro笔记本上SentenceTransformer模型的编码速度?

加速M1 Pro上BERT文本编码的方法

你在M1 Pro的Mac上使用SentenceTransformer的distilbert-base-nli-mean-tokens模型编码20newsgroups全量数据时速度极慢,可通过以下方法优化:

  • 启用Apple Silicon GPU加速
    M1 Pro的MPS后端能大幅提升张量运算效率,默认可能未开启,手动设置设备:

    import torch
    device = torch.device("mps" if torch.backends.mps.is_available() else "cpu")
    model = SentenceTransformer('distilbert-base-nli-mean-tokens').to(device)
    
  • 增大批量大小
    调整encode方法的batch_size参数,提升GPU利用率,可根据内存情况逐步测试合适的值:

    embeddings = model.encode(data, show_progress_bar=True, batch_size=64)
    
  • 截断过长文本
    BERT类模型有512 tokens的最大输入限制,提前截断文本可减少无效运算:

    from transformers import AutoTokenizer
    tokenizer = AutoTokenizer.from_pretrained('distilbert-base-nli-mean-tokens')
    def truncate_text(text):
        tokens = tokenizer.tokenize(text)
        if len(tokens) > 510:
            tokens = tokens[:510]
        return tokenizer.convert_tokens_to_string(tokens)
    data_truncated = [truncate_text(text) for text in data]
    embeddings = model.encode(data_truncated, show_progress_bar=True, device=device, batch_size=64)
    
  • 更换轻量模型
    若对精度要求不高,可使用参数更少的模型,比如all-MiniLM-L6-v2,编码速度更快:

    model = SentenceTransformer('all-MiniLM-L6-v2').to(device)
    
  • 开启多进程CPU辅助
    若GPU仍有瓶颈,可启用多进程编码,num_workers设置为CPU核心数的一半左右即可:

    embeddings = model.encode(data, show_progress_bar=True, device=device, batch_size=64, use_multiprocessing=True, num_workers=4)
    

内容的提问来源于stack exchange,提问作者Toly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 00:12:31