You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Chroma DB搭配LlamaIndex查询自定义PDF时返回None问题

问题排查与修复方案

以下是导致查询返回None的关键问题及对应修复:

1. LLMPredictor初始化错误

在CreatingChromaDB函数中,错误地将LLMPredictor实例作为参数传入自身构造函数,导致LLM无法正常工作。

修复前代码:

llm_predictor = LLMPredictor(llm=llm)
llm_predictor = LLMPredictor(llm = llm_predictor)  # 错误行

修复后代码:

llm_predictor = LLMPredictor(llm=llm)
# 移除错误的重复初始化行

2. 错误使用聊天模型对应的LLM类

部署名称为gpt-3.5-turbo(聊天模型),但使用了适用于补全模型的AzureOpenAI类,导致LLM调用失败。需替换为AzureChatOpenAI,并使用ChatLLMPredictor适配聊天模型。

修复前代码:

from langchain.llms import AzureOpenAI
llm = AzureOpenAI(deployment_name=deployment_name,openai_api_version=openai_api_version)
llm_predictor = LLMPredictor(llm=llm)

修复后代码:

from langchain.chat_models import AzureChatOpenAI  # 导入正确的聊天模型类
from llama_index import ChatLLMPredictor  # 导入聊天模型预测器

llm = AzureChatOpenAI(deployment_name=deployment_name, openai_api_version=openai_api_version)
llm_predictor = ChatLLMPredictor(llm=llm)  # 使用ChatLLMPredictor适配聊天模型

3. 加载索引时缺少ServiceContext

LoadFromDisk函数中创建索引时未传入包含嵌入模型和LLM预测器的ServiceContext,导致索引无法生成查询嵌入或调用LLM生成回答。

修复后完整LoadFromDisk函数:

from chromadb.config import Settings  # 新增导入

def LoadFromDisk(collection_name,persist_directory):
    # 重建与创建索引时一致的ServiceContext
    deployment_name = "gpt-3.5-turbo"
    openai_api_version="2023-08-30"  # 修正为Azure标准日期格式
    
    llm = AzureChatOpenAI(deployment_name=deployment_name, openai_api_version=openai_api_version)
    llm_predictor = ChatLLMPredictor(llm=llm)
    embedding_llm = LangchainEmbedding(OpenAIEmbeddings())
    
    max_input_size = 3000
    num_output = 256
    chunk_size_limit = 1000
    max_chunk_overlap = 20
    prompt_helper = PromptHelper(max_input_size=max_input_size, num_output=num_output,
                              max_chunk_overlap=max_chunk_overlap, chunk_size_limit=chunk_size_limit)
    
    service_context = ServiceContext.from_defaults(llm_predictor=llm_predictor, embed_model=embedding_llm, prompt_helper=prompt_helper)
    
    chroma_client = chromadb.Client(Settings(
        chroma_db_impl="duckdb+parquet",
        persist_directory= persist_directory))

    print(collection_name)

    # 加载集合时指定嵌入函数
    chroma_collection = chroma_client.get_or_create_collection(collection_name, embedding_function=embedding_llm)
    vector_store = ChromaVectorStore(chroma_collection=chroma_collection)
    # 传入service_context
    index = GPTVectorStoreIndex.from_vector_store(vector_store=vector_store, service_context=service_context)
    return index

4. Chroma集合加载时未指定嵌入函数

加载Chroma集合时未传入嵌入函数,导致无法生成查询嵌入以匹配文档,进而无相关上下文传入LLM,返回None。

修复点:
在LoadFromDisk中创建集合时添加embedding_function=embedding_llm参数,如上述代码所示。

5. 提示模板格式错误

模板中存在多余的双引号,导致提示解析异常。

修正后的完整模板:

TEMPLATE_STR = """Create a final answer to the given questions using the provided document excerpts(in no particular order) as references. ALWAYS include a "SOURCES" section in your answer including only the minimal set of sources needed to answer the question. Always include the Source Preview of source. If answer has step in document please response in step. If you are unable to answer the question, simply state that you do not know. Do not attempt to fabricate an answer and leave the SOURCES section empty.

---------------------
{context_str}
---------------------
Given this information, please answer the question: {query_str}
"""

6. Azure API版本格式错误

原代码中openai_api_version="30/08/2023"不符合Azure API版本格式要求,需修正为YYYY-MM-DD格式。

修复后:

openai_api_version="2023-08-30"

7. 缺少Chroma Settings导入

在vectordb.py顶部添加导入:

from chromadb.config import Settings

内容的提问来源于stack exchange,提问作者Pmd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 00:27:09