使用Hugging Face本地模型时llama-index仍要求OpenAI密钥的问题
问题描述
基于llama-index开发文档问答应用,此前使用OpenAI可正常运行。现在希望完全不依赖外部API,参照官方Hugging Face示例配置了本地LLM(StableLM),但运行代码时仍报错:
ValueError: No API key found for OpenAI.
Please set either the OPENAI_API_KEY environment variable or openai.api_key prior to initialization.
API keys can be found or created at https://platform.openai.com/account/api-keys
问题原因
你确实忽略了关键配置:llama-index默认使用OpenAI的嵌入模型(Embedding),即使替换了本地LLM,文档向量化阶段还是会调用OpenAI API,因此触发了API key错误。官方文档提到的“设置本地嵌入模型”正是解决这个问题的核心。
解决方案
需要添加本地嵌入模型到ServiceContext中,具体步骤如下:
- 导入
HuggingFaceEmbedding模块 - 创建本地嵌入模型实例
- 将嵌入模型传入
ServiceContext.from_defaults方法
修改后的关键代码片段
# 新增导入语句 from llama_index.embeddings import HuggingFaceEmbedding def construct_index(directory_path): # ... 保留原有代码 ... # 初始化本地嵌入模型 embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-mpnet-base-v2") # 将嵌入模型加入ServiceContext配置 service_context = ServiceContext.from_defaults( chunk_size=1024, llm=llm, embed_model=embed_model # 新增该行 ) # ... 保留原有代码 ...
完整修改后的construct_index函数
def construct_index(directory_path): max_input_size = 4096 num_outputs = 512 chunk_overlap_ratio = 0.1 chunk_size_limit = 600 system_prompt = """<|SYSTEM|># StableLM Tuned (Alpha version) - StableLM is a helpful and harmless open-source AI language model developed by StabilityAI. - StableLM is excited to be able to help the user, but will refuse to do anything that could be considered harmful to the user. - StableLM is more than just an information source, StableLM is also able to write poetry, short stories, and make jokes. - StableLM will refuse to participate in anything that could harm a human. """ query_wrapper_prompt = SimpleInputPrompt("<|USER|>{query_str}<|ASSISTANT|>") llm = HuggingFaceLLM( context_window=4096, max_new_tokens=256, generate_kwargs={"temperature": 0.7, "do_sample": False}, system_prompt=system_prompt, query_wrapper_prompt=query_wrapper_prompt, tokenizer_name="StabilityAI/stablelm-tuned-alpha-3b", model_name="StabilityAI/stablelm-tuned-alpha-3b", device_map="auto", stopping_ids=[50278, 50279, 50277, 1, 0], tokenizer_kwargs={"max_length": 4096}, # model_kwargs={"torch_dtype": torch.float16} # 若使用CUDA可取消注释以节省内存 ) # 新增:配置本地嵌入模型 embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-mpnet-base-v2") service_context = ServiceContext.from_defaults( chunk_size=1024, llm=llm, embed_model=embed_model ) documents = SimpleDirectoryReader(directory_path).load_data() index = VectorStoreIndex.from_documents(documents, service_context=service_context) index.storage_context.persist(persist_dir=storage_path) return index
修改后,文档向量化、LLM推理全流程都会在本地运行,不再依赖OpenAI API。如果硬件资源有限,可以选择更轻量的嵌入模型,比如sentence-transformers/all-MiniLM-L6-v2。
内容的提问来源于stack exchange,提问作者Mikey A. Leonetti
相关产品推荐
相关产品推荐

