You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python与LangChain拆分文档块时遇类型错误的原因排查

问题解决:LangChain文本拆分触发TypeError错误

错误原因

你的代码中,parser.from_file('..\\mql5_new_compressed.pdf')["content"]已经将PDF解析后的文本内容提取为字符串类型,但后续调用text_splitter.create_documents([text["content"]])时,试图对字符串text使用字符串索引["content"],这就触发了TypeError: string indices must be integers, not 'str'错误。

修正步骤

  • 直接使用已提取的字符串变量text作为拆分输入,无需再取["content"]
  • 增加空值校验,避免PDF解析内容为空时引发异常
  • 补充向量库存储的完整逻辑,完成从文本拆分到存入Chroma的流程

修正后的完整代码

# Description: This script is used to extract text from a PDF file and store it in a Chroma vector database.
import tika
from tika import parser
from langchain.text_splitter import CharacterTextSplitter
from langchain_community.embeddings import HuggingFaceEmbeddings
from langchain_community.vectorstores import Chroma

# 初始化tika服务
tika.initVM()

# 解析PDF并提取文本
parsed_pdf = parser.from_file('..\\mql5_new_compressed.pdf')
text = parsed_pdf.get("content", "")

# 校验文本有效性
if not text.strip():
    raise ValueError("PDF文件中未提取到有效文本内容")

# 初始化文本拆分器
text_splitter = CharacterTextSplitter(chunk_size=1000, chunk_overlap=200)

# 拆分文本为块
chunks = text_splitter.create_documents([text])

# 初始化嵌入模型
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")

# 将文本块存入Chroma向量库
db = Chroma.from_documents(chunks, embeddings, persist_directory="./chroma_db")
db.persist()

print("文本拆分与向量库存储完成")

关键说明

  • parser.from_file()返回字典结构,["content"]对应解析出的文本字符串,赋值给text后,text已是字符串,不能再用text["content"]索引
  • 加入tika.initVM()确保tika服务正常初始化(部分运行环境必需)
  • 使用get("content", "")避免字典无content键时触发KeyError
  • 补充了向量库存储代码,完成检索数据库的搭建流程

内容的提问来源于stack exchange,提问作者Christian Bannard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 08:22:50