You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用HuggingFace与LangChain创建4bit量化Llama3模型Embedding报错求助

问题修正方案

你遇到的错误本质是选错了模型类型:你用的unsloth/llama-3-8b-Instruct-bnb-4bit是一个4bit量化的对话大语言模型(LLM),并非专门的文本嵌入模型,而且bitsandbytes量化的模型无法直接被HuggingFaceEmbeddings加载——它需要的是支持嵌入任务的模型(比如Sentence-BERT系列),而非量化的LLM。

正确的解决办法

方案1:使用专门的文本嵌入模型(推荐)

替换为Sentence-Transformers系列的嵌入模型,这类模型专门为文本嵌入任务优化,轻量且高效,示例代码如下:

from langchain_huggingface import HuggingFaceEmbeddings
from langchain_community.vectorstores import Chroma
import torch

# 更换为专门的嵌入模型
model_name = "sentence-transformers/all-MiniLM-L6-v2"
model_kwargs = {'device': 'cuda' if torch.cuda.is_available() else 'cpu'}
encode_kwargs = {'normalize_embeddings': True}  # 嵌入归一化通常有助于提升检索效果

hf = HuggingFaceEmbeddings(
    model_name=model_name,
    model_kwargs=model_kwargs,
    encode_kwargs=encode_kwargs
)

# 后续创建向量数据库的代码保持不变
vectorstore = Chroma.from_documents(documents=splits, embedding=hf)

方案2:若坚持使用LLM生成嵌入(不推荐)

如果一定要用LLM来生成嵌入,需要使用非量化版本的模型,并确保模型支持嵌入任务,同时添加必要的参数:

from langchain_huggingface import HuggingFaceEmbeddings
from langchain_community.vectorstores import Chroma
import torch

# 使用非量化的LLM模型
model_name = "unsloth/llama-3-8b-Instruct"
model_kwargs = {
    'device': 'cuda' if torch.cuda.is_available() else 'cpu',
    'trust_remote_code': True,  # 加载自定义模型需要开启该参数
    'task_type': 'embeddings'   # 指定模型用于嵌入任务
}
encode_kwargs = {'normalize_embeddings': False}

hf = HuggingFaceEmbeddings(
    model_name=model_name,
    model_kwargs=model_kwargs,
    encode_kwargs=encode_kwargs
)

vectorstore = Chroma.from_documents(documents=splits, embedding=hf)

注意:LLM生成嵌入的效率远低于专门的嵌入模型,且效果不一定更好,仅在特殊场景下考虑使用。

内容的提问来源于stack exchange,提问作者brian chow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:12:37