如何在LangChain的FAISS向量库中为数据集添加Semantic Chunking?
解决LangChain中SemanticChunker结合FAISS的实现问题
你遇到的问题大概率是数据结构不匹配导致的:SemanticChunker处理后输出的是LangChain的Document对象列表,但你后续仍用FAISS.from_texts()(该方法需要纯字符串列表),自然会报错。下面分两种原始数据格式给出正确实现:
情况1:原始docs是纯文本字符串列表
如果你的docs是["文本内容1", "文本内容2", ...]这种纯字符串列表,正确流程是用SemanticChunker.create_documents()生成Document对象,再用FAISS.from_documents()存入向量库:
from langchain_community.vectorstores import FAISS from langchain_openai import OpenAIEmbeddings from langchain_experimental.text_splitter import SemanticChunker # 原始纯文本列表 raw_docs = ["这里是你的大段文本内容1", "这里是你的大段文本内容2"] # 初始化语义分块器 embeddings = OpenAIEmbeddings(openai_api_key=OPENAI_API_KEY) text_splitter = SemanticChunker(embeddings) # 生成分块后的Document对象列表 split_docs = text_splitter.create_documents(raw_docs) # 存入FAISS向量库(注意用from_documents而非from_texts) vectorstore = FAISS.from_documents(split_docs, embedding=embeddings) retriever = vectorstore.as_retriever()
情况2:原始docs是LangChain Document对象列表
如果你的docs已经是Document(page_content="...", metadata={...})格式的列表,直接用SemanticChunker.split_documents()处理:
from langchain_community.vectorstores import FAISS from langchain_openai import OpenAIEmbeddings from langchain_experimental.text_splitter import SemanticChunker from langchain_core.documents import Document # 原始Document对象列表 raw_docs = [ Document(page_content="这里是你的大段文本内容1", metadata={"source": "doc1"}), Document(page_content="这里是你的大段文本内容2", metadata={"source": "doc2"}) ] # 初始化语义分块器 embeddings = OpenAIEmbeddings(openai_api_key=OPENAI_API_KEY) text_splitter = SemanticChunker(embeddings) # 对已有Document进行语义分块 split_docs = text_splitter.split_documents(raw_docs) # 存入FAISS向量库 vectorstore = FAISS.from_documents(split_docs, embedding=embeddings) retriever = vectorstore.as_retriever()
关键注意点
- 语义分块后必须用
FAISS.from_documents(),而不是from_texts(),因为分块结果是带元数据的Document对象; - SemanticChunker依赖嵌入模型,确保你的OpenAI API密钥有效且网络正常;
- 如果分块效果不理想,可以调整SemanticChunker的参数,比如
breakpoint_threshold_type(可选"percentile"或"standard_deviation")来控制分块粒度。
内容的提问来源于stack exchange,提问作者user17811469
相关产品推荐
相关产品推荐

