如何加速M1 Pro笔记本上SentenceTransformer模型的编码速度?
加速M1 Pro上BERT文本编码的方法
你在M1 Pro的Mac上使用SentenceTransformer的distilbert-base-nli-mean-tokens模型编码20newsgroups全量数据时速度极慢,可通过以下方法优化:
启用Apple Silicon GPU加速
M1 Pro的MPS后端能大幅提升张量运算效率,默认可能未开启,手动设置设备:import torch device = torch.device("mps" if torch.backends.mps.is_available() else "cpu") model = SentenceTransformer('distilbert-base-nli-mean-tokens').to(device)增大批量大小
调整encode方法的batch_size参数,提升GPU利用率,可根据内存情况逐步测试合适的值:embeddings = model.encode(data, show_progress_bar=True, batch_size=64)截断过长文本
BERT类模型有512 tokens的最大输入限制,提前截断文本可减少无效运算:from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained('distilbert-base-nli-mean-tokens') def truncate_text(text): tokens = tokenizer.tokenize(text) if len(tokens) > 510: tokens = tokens[:510] return tokenizer.convert_tokens_to_string(tokens) data_truncated = [truncate_text(text) for text in data] embeddings = model.encode(data_truncated, show_progress_bar=True, device=device, batch_size=64)更换轻量模型
若对精度要求不高,可使用参数更少的模型,比如all-MiniLM-L6-v2,编码速度更快:model = SentenceTransformer('all-MiniLM-L6-v2').to(device)开启多进程CPU辅助
若GPU仍有瓶颈,可启用多进程编码,num_workers设置为CPU核心数的一半左右即可:embeddings = model.encode(data, show_progress_bar=True, device=device, batch_size=64, use_multiprocessing=True, num_workers=4)
内容的提问来源于stack exchange,提问作者Toly
相关产品推荐
相关产品推荐

