使用Python与LangChain拆分文档块时遇类型错误的原因排查
问题解决:LangChain文本拆分触发TypeError错误
错误原因
你的代码中,parser.from_file('..\\mql5_new_compressed.pdf')["content"]已经将PDF解析后的文本内容提取为字符串类型,但后续调用text_splitter.create_documents([text["content"]])时,试图对字符串text使用字符串索引["content"],这就触发了TypeError: string indices must be integers, not 'str'错误。
修正步骤
- 直接使用已提取的字符串变量
text作为拆分输入,无需再取["content"] - 增加空值校验,避免PDF解析内容为空时引发异常
- 补充向量库存储的完整逻辑,完成从文本拆分到存入Chroma的流程
修正后的完整代码
# Description: This script is used to extract text from a PDF file and store it in a Chroma vector database. import tika from tika import parser from langchain.text_splitter import CharacterTextSplitter from langchain_community.embeddings import HuggingFaceEmbeddings from langchain_community.vectorstores import Chroma # 初始化tika服务 tika.initVM() # 解析PDF并提取文本 parsed_pdf = parser.from_file('..\\mql5_new_compressed.pdf') text = parsed_pdf.get("content", "") # 校验文本有效性 if not text.strip(): raise ValueError("PDF文件中未提取到有效文本内容") # 初始化文本拆分器 text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=200) # 拆分文本为块 chunks = text_splitter.create_documents([text]) # 初始化嵌入模型 embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2") # 将文本块存入Chroma向量库 db = Chroma.from_documents(chunks, embeddings, persist_directory="./chroma_db") db.persist() print("文本拆分与向量库存储完成")
关键说明
parser.from_file()返回字典结构,["content"]对应解析出的文本字符串,赋值给text后,text已是字符串,不能再用text["content"]索引- 加入
tika.initVM()确保tika服务正常初始化(部分运行环境必需) - 使用
get("content", "")避免字典无content键时触发KeyError - 补充了向量库存储代码,完成检索数据库的搭建流程
内容的提问来源于stack exchange,提问作者Christian Bannard
相关产品推荐
相关产品推荐

