点击Streamlit按钮触发Chroma ValueError:Expected IDs非空列表排查
问题:点击"Continue to Questionnaire"或"Chat with BotAI"按钮触发Chroma ValueError错误
错误详情
**ValueError:** Expected IDs to be a non-empty list, got [] **Traceback:** File "C:\Users\scite\Desktop\HAMBOTAI\HAMBotAI\HAMBotAI\homehambotai.py", line 96, in app db = Chroma.from_documents(texts, embeddings) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\langchain_community\vectorstores\chroma.py", line 771, in from_documents return cls.from_texts( ^^^^^^^^^^^^^^^ File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\langchain_community\vectorstores\chroma.py", line 729, in from_texts chroma_collection.add_texts( File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\langchain_community\vectorstores\chroma.py", line 324, in add_texts self._collection.upsert( File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\chromadb\api\models\Collection.py", line 449, in upsert ) = self._validate_embedding_set( ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\chromadb\api\models\Collection.py", line 512, in _validate_embedding_set valid_ids = validate_ids(maybe_cast_one_to_many_ids(ids)) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\scite\AppData\Roaming\Python\Python311\site-packages\chromadb\api\types.py", line 228, in validate_ids raise ValueError(f"Expected IDs to be a non-empty list, got {ids}")
相关代码片段
if 'processed' in query_params: # Create a temporary text file with tempfile.NamedTemporaryFile(mode="w", delete=False, suffix=".txt") as temp_file: temp_file.write(text) temp_file_path = temp_file.name # load document loader = TextLoader(temp_file_path) documents = loader.load() # split the documents into chunks text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0) texts = text_splitter.split_documents(documents) # select which embeddings we want to use embeddings = OpenAIEmbeddings() # ids =[str(i) for i in range(1, len(texts) + 1)] # create the vectorestore to use as the index db = Chroma.from_documents(texts, embeddings) # expose this index in a retriever interface retriever = db.as_retriever(search_type="similarity", search_kwargs={"k": 2}) # create a chain to answer questions qa = ConversationalRetrievalChain.from_llm(OpenAI(), retriever) chat_history = [] # query = "What's the Name of patient and doctor as mentioned in the data?" # result = qa({"question": query, "chat_history": chat_history}) # st.write("Patient and Doctor name:", result['answer']) # # chat_history = [(query, result["answer"])] query = "Provide summary of medical and health related info from this data in points, every point should be in new line (Formatted in HTML)?" result = qa({"question": query, "chat_history": chat_history}) toshow = result['answer'] # chat_history = [(query, result["answer"])] # chat_history.append((query, result["answer"])) # print(chat_history) st.title("Data Fetched From Your Health & Medical Reports") components.html( f""" {toshow} """, height=250, scrolling=True, ) if st.button('Continue to Questionarrie'): st.write('Loading') st.text("(OR)") if st.button('Chat with BotAI'): st.title("Chat with BotAI")
核心原因
Streamlit的按钮点击会触发整个脚本重新运行:
- 点击按钮时,URL的
query_params仍保留processed标识,代码会再次进入if 'processed' in query_params分支 - 此时
text变量已丢失有效内容(没有重新生成/上传文本),导致加载的文档为空,拆分后的texts是空列表 - Chroma调用
from_documents时,空的texts会生成空的ID列表,触发校验错误
解决方案
1. 用Streamlit会话状态缓存处理结果
把向量库、QA链、摘要结果等存入st.session_state,避免重复执行文档处理逻辑:
if 'processed' in query_params: # Create a temporary text file with tempfile.NamedTemporaryFile(mode="w", delete=False, suffix=".txt") as temp_file: temp_file.write(text) temp_file_path = temp_file.name # load document loader = TextLoader(temp_file_path) documents = loader.load() # split the documents into chunks text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=0) texts = text_splitter.split_documents(documents) # 新增:校验texts是否为空 if not texts: st.error("上传的文本内容为空,无法生成向量库") else: embeddings = OpenAIEmbeddings() db = Chroma.from_documents(texts, embeddings) retriever = db.as_retriever(search_type="similarity", search_kwargs={"k": 2}) qa = ConversationalRetrievalChain.from_llm(OpenAI(), retriever) chat_history = [] query = "Provide summary of medical and health related info from this data in points, every point should be in new line (Formatted in HTML)?" result = qa({"question": query, "chat_history": chat_history}) toshow = result['answer'] # 缓存结果到会话状态 st.session_state['toshow'] = toshow st.session_state['qa_chain'] = qa # 从会话状态读取结果,避免重复处理 if 'toshow' in st.session_state: st.title("Data Fetched From Your Health & Medical Reports") components.html( f""" {st.session_state['toshow']} """, height=250, scrolling=True, ) if st.button('Continue to Questionarrie'): st.write('Loading') # 可在这里切换会话状态,比如st.session_state['page'] = 'questionnaire' st.text("(OR)") if st.button('Chat with BotAI'): st.title("Chat with BotAI") # 可使用缓存的qa_chain进行对话
2. 按钮点击后更新页面状态
点击按钮后,通过会话状态标记当前页面,避免重新进入processed分支:
if st.button('Continue to Questionarrie'): st.session_state['current_page'] = 'questionnaire' st.experimental_rerun() # 在脚本开头判断页面状态 if 'current_page' in st.session_state and st.session_state['current_page'] == 'questionnaire': # 渲染问卷页面逻辑 st.write("问卷页面") else: # 原有的文档处理和结果展示逻辑 pass
内容的提问来源于stack exchange,提问作者SCITECHE
相关产品推荐
相关产品推荐

